{
  "id": 158202,
  "title": "Error when doing k-fold training on TPU.",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/158202",
  "author_name": "Urvish",
  "post_date": "2020-06-13T12:20:53.067000",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I am trying to use <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">this</a> public kernel and modify it to do k-fold training on TPU. But I am getting the following error after first fold finishes. Thanks for the kernel <a href=\"/riblidezso\">@riblidezso</a> .</p>\n\n<p><code>\nNotFoundError: {{function_node __inference_distributed_train_step_139198}} Resource worker/Adam/iter/replica_6_57590/N10tensorflow3VarE does not exist.\n     [[{{node cluster_distributed_train_step/_execute_6_0}}]]\n</code>\nI made changes like calling <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> before each of the fold. I also added <code>drop_reminder=True</code> where needed. But I am getting above error when starting the second fold. Not sure why yet. Seems like there is an issue in creating a distributed dataset in the second fold. So how can I fix this? I am not able to paste the full traceback somehow here not sure why. </p>",
  "messages": [
    {
      "id": 885144,
      "postDate": "2020-06-14T00:14:47.010Z",
      "content": "<p>My dumb fix is to have K-kernels for K-folds (as a last resort :D )</p>",
      "rawMarkdown": "My dumb fix is to have K-kernels for K-folds (as a last resort :D )",
      "votes": 3,
      "replies": [
        {
          "id": 885259,
          "postDate": "2020-06-14T04:14:23.053Z",
          "content": "<p>yeah but that is not going to help much as now I am not going to have enough time to do so. I found some kernels from <code>Flower classification challenge on TPU</code> which might be helpful.</p>",
          "rawMarkdown": "yeah but that is not going to help much as now I am not going to have enough time to do so. I found some kernels from `Flower classification challenge on TPU` which might be helpful.",
          "votes": 1
        }
      ]
    },
    {
      "id": 885119,
      "postDate": "2020-06-13T22:48:28.813Z",
      "content": "<p>Just simply don't do this, as this causes the problem of the deleted variable.\n<code>\ntf.tpu.experimental.initialize_tpu_system(tpu)\n</code>\nthis is a brute force way to fix issues with Keras taking and leaving around too much memory.</p>\n\n<p>With my kernel you can have repeated train/predict calls (10+) without any memory issue.</p>",
      "rawMarkdown": "Just simply don't do this, as this causes the problem of the deleted variable.\n```\ntf.tpu.experimental.initialize_tpu_system(tpu)\n```\nthis is a brute force way to fix issues with Keras taking and leaving around too much memory.\n\nWith my kernel you can have repeated train/predict calls (10+) without any memory issue.",
      "votes": 3,
      "replies": [
        {
          "id": 885260,
          "postDate": "2020-06-14T04:17:01.680Z",
          "content": "<p>I did not get the first point <code>Just do not do it</code>, what do you mean? Your kernel is great and I think I will make it work for me for k-folds. I used the above-suggested way but then the data is not getting distributed somehow even if I call the function. Not sure why. </p>",
          "rawMarkdown": "I did not get the first point `Just do not do it`, what do you mean? Your kernel is great and I think I will make it work for me for k-folds. I used the above-suggested way but then the data is not getting distributed somehow even if I call the function. Not sure why. "
        },
        {
          "id": 885339,
          "postDate": "2020-06-14T05:43:32.103Z",
          "content": "<p>I think I misunderstood your first point. You want to say that <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> do not use this at the first line of each fold? Am I right?</p>",
          "rawMarkdown": "I think I misunderstood your first point. You want to say that `tf.tpu.experimental.initialize_tpu_system(tpu)` do not use this at the first line of each fold? Am I right?"
        },
        {
          "id": 885360,
          "postDate": "2020-06-14T06:07:43.830Z",
          "content": "<p>I removed this line from the beginning of the fold but then got the resource exhausted error as mentioned in other discussions. I am doing following steps while training \n1. Re-initialize the strategy, tpu and global batch size.\n2. Make the distributed data.\n3. Load losses and metrics\n4. Load model and optimizer \n5.  Train and fine-tune model\n6. Predict on test data.\nNow if I remove <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code>, then I got resource exhausted error and if not then above not found error. What steps should I remove from the loop in order to make it work? </p>",
          "rawMarkdown": "I removed this line from the beginning of the fold but then got the resource exhausted error as mentioned in other discussions. I am doing following steps while training \n1. Re-initialize the strategy, tpu and global batch size.\n2. Make the distributed data.\n3. Load losses and metrics\n4. Load model and optimizer \n5.  Train and fine-tune model\n6. Predict on test data.\nNow if I remove `tf.tpu.experimental.initialize_tpu_system(tpu)`, then I got resource exhausted error and if not then above not found error. What steps should I remove from the loop in order to make it work? "
        },
        {
          "id": 885782,
          "postDate": "2020-06-14T13:36:34.953Z",
          "content": "<p>Sorry If I was not totally clear. Yes I meant do no reinitizalize the TPU.</p>\n\n<p>You get resource exhausted error because you reinitialize stuff for each fold. You don't need to reinitialize the TPU or the model or optimizers. </p>\n\n<p>Just reset the model weights, and redefine the training and validation datasets.</p>\n\n<p>My pseudo code for cross validation:</p>\n\n<p>```\nfor tr_idx, test_idx in KFold(K).split(X):\n    ### Splitting data ### \n    ...</p>\n\n<pre><code>### Creating Datasets ###\ntrain_dist_dataset = create_dist_dataset(X_train, y_train, training = True)\nval_dist_dataset = create_dist_dataset(X_val)\n\nmodel.set_weights(initial_w)\ntrain()\npredict()\n</code></pre>\n\n<p>```</p>\n\n<p>I hope this helps. Good luck.</p>",
          "rawMarkdown": "Sorry If I was not totally clear. Yes I meant do no reinitizalize the TPU.\n\nYou get resource exhausted error because you reinitialize stuff for each fold. You don't need to reinitialize the TPU or the model or optimizers. \n\nJust reset the model weights, and redefine the training and validation datasets.\n\nMy pseudo code for cross validation:\n\n```\nfor tr_idx, test_idx in KFold(K).split(X):\n    ### Splitting data ### \n    ...\n\n    ### Creating Datasets ###\n    train_dist_dataset = create_dist_dataset(X_train, y_train, training = True)\n    val_dist_dataset = create_dist_dataset(X_val)\n\n    model.set_weights(initial_w)\n    train()\n    predict()\n```\n  \nI hope this helps. Good luck.\n",
          "votes": 3
        },
        {
          "id": 886728,
          "postDate": "2020-06-15T08:38:35.537Z",
          "content": "<p>Thanks for the help. I got your point. I will do this quickly and will see if I can get more boost in score. Thanks.</p>",
          "rawMarkdown": "Thanks for the help. I got your point. I will do this quickly and will see if I can get more boost in score. Thanks."
        }
      ]
    },
    {
      "id": 885384,
      "postDate": "2020-06-14T06:45:11.950Z",
      "content": "<p>I solved it finally. I realized how much silly the mistake was. The fix was that I had to define all the functions inside the for loop which iterates over each fold. For example, define the load model, create a dataset, train step etc inside the loop and it worked. Some useful resources were <a href=\"https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu#Optimized-custom-training-loop\">this</a> and <a href=\"https://www.kaggle.com/dimitreoliveira/flower-with-tpus-k-fold-optimized-training-loops\">this</a></p>",
      "rawMarkdown": "I solved it finally. I realized how much silly the mistake was. The fix was that I had to define all the functions inside the for loop which iterates over each fold. For example, define the load model, create a dataset, train step etc inside the loop and it worked. Some useful resources were [this](https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu#Optimized-custom-training-loop) and [this](https://www.kaggle.com/dimitreoliveira/flower-with-tpus-k-fold-optimized-training-loops)",
      "votes": 1
    },
    {
      "id": 885114,
      "postDate": "2020-06-13T22:25:08.047Z",
      "content": "<p>Beware of 5GB limit for output. You can save only 2 models as one is 2+gb</p>",
      "rawMarkdown": "Beware of 5GB limit for output. You can save only 2 models as one is 2+gb",
      "votes": 1,
      "replies": [
        {
          "id": 885261,
          "postDate": "2020-06-14T04:17:52.097Z",
          "content": "<p>I do not want to save all models, at this stage I need better predictions only so that I can later make ensemble out of it. </p>",
          "rawMarkdown": "I do not want to save all models, at this stage I need better predictions only so that I can later make ensemble out of it. "
        }
      ]
    },
    {
      "id": 884497,
      "postDate": "2020-06-13T12:20:53.067Z",
      "content": "<p>I am trying to use <a href=\"https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large\">this</a> public kernel and modify it to do k-fold training on TPU. But I am getting the following error after first fold finishes. Thanks for the kernel <a href=\"/riblidezso\">@riblidezso</a> .</p>\n\n<p><code>\nNotFoundError: {{function_node __inference_distributed_train_step_139198}} Resource worker/Adam/iter/replica_6_57590/N10tensorflow3VarE does not exist.\n     [[{{node cluster_distributed_train_step/_execute_6_0}}]]\n</code>\nI made changes like calling <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> before each of the fold. I also added <code>drop_reminder=True</code> where needed. But I am getting above error when starting the second fold. Not sure why yet. Seems like there is an issue in creating a distributed dataset in the second fold. So how can I fix this? I am not able to paste the full traceback somehow here not sure why. </p>",
      "rawMarkdown": "I am trying to use [this](https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large) public kernel and modify it to do k-fold training on TPU. But I am getting the following error after first fold finishes. Thanks for the kernel @riblidezso .\n\n```\nNotFoundError: {{function_node __inference_distributed_train_step_139198}} Resource worker/Adam/iter/replica_6_57590/N10tensorflow3VarE does not exist.\n\t [[{{node cluster_distributed_train_step/_execute_6_0}}]]\n```\nI made changes like calling `tf.tpu.experimental.initialize_tpu_system(tpu)` before each of the fold. I also added `drop_reminder=True` where needed. But I am getting above error when starting the second fold. Not sure why yet. Seems like there is an issue in creating a distributed dataset in the second fold. So how can I fix this? I am not able to paste the full traceback somehow here not sure why. ",
      "votes": 2
    },
    {
      "id": 884509,
      "postDate": "2020-06-13T12:25:28.017Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 885144,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2020-06-14T00:14:47.010000",
      "content": "<p>My dumb fix is to have K-kernels for K-folds (as a last resort :D )</p>",
      "votes": 3,
      "replies": [
        {
          "id": 885259,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-06-14T04:14:23.053000",
          "content": "<p>yeah but that is not going to help much as now I am not going to have enough time to do so. I found some kernels from <code>Flower classification challenge on TPU</code> which might be helpful.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 885119,
      "author_name": "Dezső Ribli",
      "author_url": "",
      "post_date": "2020-06-13T22:48:28.813000",
      "content": "<p>Just simply don't do this, as this causes the problem of the deleted variable.\n<code>\ntf.tpu.experimental.initialize_tpu_system(tpu)\n</code>\nthis is a brute force way to fix issues with Keras taking and leaving around too much memory.</p>\n\n<p>With my kernel you can have repeated train/predict calls (10+) without any memory issue.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 885260,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-06-14T04:17:01.680000",
          "content": "<p>I did not get the first point <code>Just do not do it</code>, what do you mean? Your kernel is great and I think I will make it work for me for k-folds. I used the above-suggested way but then the data is not getting distributed somehow even if I call the function. Not sure why. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885339,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-06-14T05:43:32.103000",
          "content": "<p>I think I misunderstood your first point. You want to say that <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code> do not use this at the first line of each fold? Am I right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885360,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-06-14T06:07:43.830000",
          "content": "<p>I removed this line from the beginning of the fold but then got the resource exhausted error as mentioned in other discussions. I am doing following steps while training \n1. Re-initialize the strategy, tpu and global batch size.\n2. Make the distributed data.\n3. Load losses and metrics\n4. Load model and optimizer \n5.  Train and fine-tune model\n6. Predict on test data.\nNow if I remove <code>tf.tpu.experimental.initialize_tpu_system(tpu)</code>, then I got resource exhausted error and if not then above not found error. What steps should I remove from the loop in order to make it work? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885782,
          "author_name": "Dezső Ribli",
          "author_url": "",
          "post_date": "2020-06-14T13:36:34.953000",
          "content": "<p>Sorry If I was not totally clear. Yes I meant do no reinitizalize the TPU.</p>\n\n<p>You get resource exhausted error because you reinitialize stuff for each fold. You don't need to reinitialize the TPU or the model or optimizers. </p>\n\n<p>Just reset the model weights, and redefine the training and validation datasets.</p>\n\n<p>My pseudo code for cross validation:</p>\n\n<p>```\nfor tr_idx, test_idx in KFold(K).split(X):\n    ### Splitting data ### \n    ...</p>\n\n<pre><code>### Creating Datasets ###\ntrain_dist_dataset = create_dist_dataset(X_train, y_train, training = True)\nval_dist_dataset = create_dist_dataset(X_val)\n\nmodel.set_weights(initial_w)\ntrain()\npredict()\n</code></pre>\n\n<p>```</p>\n\n<p>I hope this helps. Good luck.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 886728,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-06-15T08:38:35.537000",
          "content": "<p>Thanks for the help. I got your point. I will do this quickly and will see if I can get more boost in score. Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 885384,
      "author_name": "Urvish",
      "author_url": "",
      "post_date": "2020-06-14T06:45:11.950000",
      "content": "<p>I solved it finally. I realized how much silly the mistake was. The fix was that I had to define all the functions inside the for loop which iterates over each fold. For example, define the load model, create a dataset, train step etc inside the loop and it worked. Some useful resources were <a href=\"https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu#Optimized-custom-training-loop\">this</a> and <a href=\"https://www.kaggle.com/dimitreoliveira/flower-with-tpus-k-fold-optimized-training-loops\">this</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 885114,
      "author_name": "Miroslav Valan",
      "author_url": "",
      "post_date": "2020-06-13T22:25:08.047000",
      "content": "<p>Beware of 5GB limit for output. You can save only 2 models as one is 2+gb</p>",
      "votes": 1,
      "replies": [
        {
          "id": 885261,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-06-14T04:17:52.097000",
          "content": "<p>I do not want to save all models, at this stage I need better predictions only so that I can later make ensemble out of it. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 884509,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-13T12:25:28.017000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "885144": "My dumb fix is to have K-kernels for K-folds (as a last resort :D )",
    "885119": "Just simply don't do this, as this causes the problem of the deleted variable.\n```\ntf.tpu.experimental.initialize_tpu_system(tpu)\n```\nthis is a brute force way to fix issues with Keras taking and leaving around too much memory.\n\nWith my kernel you can have repeated train/predict calls (10+) without any memory issue.",
    "885384": "I solved it finally. I realized how much silly the mistake was. The fix was that I had to define all the functions inside the for loop which iterates over each fold. For example, define the load model, create a dataset, train step etc inside the loop and it worked. Some useful resources were [this](https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu#Optimized-custom-training-loop) and [this](https://www.kaggle.com/dimitreoliveira/flower-with-tpus-k-fold-optimized-training-loops)",
    "885114": "Beware of 5GB limit for output. You can save only 2 models as one is 2+gb",
    "884497": "I am trying to use [this](https://www.kaggle.com/riblidezso/train-from-mlm-finetuned-xlm-roberta-large) public kernel and modify it to do k-fold training on TPU. But I am getting the following error after first fold finishes. Thanks for the kernel @riblidezso .\n\n```\nNotFoundError: {{function_node __inference_distributed_train_step_139198}} Resource worker/Adam/iter/replica_6_57590/N10tensorflow3VarE does not exist.\n\t [[{{node cluster_distributed_train_step/_execute_6_0}}]]\n```\nI made changes like calling `tf.tpu.experimental.initialize_tpu_system(tpu)` before each of the fold. I also added `drop_reminder=True` where needed. But I am getting above error when starting the second fold. Not sure why yet. Seems like there is an issue in creating a distributed dataset in the second fold. So how can I fix this? I am not able to paste the full traceback somehow here not sure why. ",
    "884509": ""
  }
}