{
  "id": 138269,
  "title": "TPU questions",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/138269",
  "author_name": "",
  "post_date": "2020-03-24T10:39:26.152572800Z",
  "votes": 14,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I think it makes sense to have a dedicated thread for some minor TPU related questions.</p>\n\n<p>I will start:</p>\n\n<p>As far as I can see, Internet access is always possible in TPU kernels. Isn't that super risky for submission kernels? I can just pull information from anywhere while subbing. Is this allowed?</p>",
  "messages": [
    {
      "id": "784575",
      "postDate": "03/24/2020 10:39:26",
      "content": "<p>I think it makes sense to have a dedicated thread for some minor TPU related questions.</p>\n\n<p>I will start:</p>\n\n<p>As far as I can see, Internet access is always possible in TPU kernels. Isn't that super risky for submission kernels? I can just pull information from anywhere while subbing. Is this allowed?</p>",
      "rawMarkdown": "I think it makes sense to have a dedicated thread for some minor TPU related questions.\n\nI will start:\n\nAs far as I can see, Internet access is always possible in TPU kernels. Isn't that super risky for submission kernels? I can just pull information from anywhere while subbing. Is this allowed?",
      "votes": null
    },
    {
      "id": "785280",
      "postDate": "03/25/2020 00:05:30",
      "content": "<p>Internet is allowed because TPUs read data from GCS (Google Cloud Storage) so your notebook needs access to that.\nExternal data is allowed in the competition but must respect the external data rules:</p>\n\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n</blockquote>",
      "rawMarkdown": "Internet is allowed because TPUs read data from GCS (Google Cloud Storage) so your notebook needs access to that.\nExternal data is allowed in the competition but must respect the external data rules:\n\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.",
      "votes": null
    },
    {
      "id": "785643",
      "postDate": "03/25/2020 08:22:07",
      "content": "<p>So using an external API to translate test data is OK? I posted it to the thread.</p>",
      "rawMarkdown": "So using an external API to translate test data is OK? I posted it to the thread.",
      "votes": null
    },
    {
      "id": "787384",
      "postDate": "03/26/2020 18:49:30",
      "content": "<p>Is there some way to use private data in TPU kernels? </p>",
      "rawMarkdown": "Is there some way to use private data in TPU kernels?",
      "votes": null
    },
    {
      "id": "787390",
      "postDate": "03/26/2020 18:52:46",
      "content": "<p>Interesting Question, I think it's possible because if we look at the Getting Started TF notebook, there's a line which calls KaggleDatasets(). something, that should automatically put your data to the nearest GCP bucket available (at least that's what the impression I ha e Collected from Martin's comments somewhere)!\nAlso you can simply use external datasets, no Psi?</p>",
      "rawMarkdown": "Interesting Question, I think it's possible because if we look at the Getting Started TF notebook, there's a line which calls KaggleDatasets(). something, that should automatically put your data to the nearest GCP bucket available (at least that's what the impression I ha e Collected from Martin's comments somewhere)!\nAlso you can simply use external datasets, no Psi?",
      "votes": null
    },
    {
      "id": "787481",
      "postDate": "03/26/2020 20:18:48",
      "content": "<p>I am debating whether it makes sense to work with TPUs in general because at some point I might need private data I want to import. Maybe encrypting the content and then decrypting in kernel?</p>\n\n<p>But you are right, as Internet access is on pulling it from external should also work.</p>",
      "rawMarkdown": "I am debating whether it makes sense to work with TPUs in general because at some point I might need private data I want to import. Maybe encrypting the content and then decrypting in kernel?\n\nBut you are right, as Internet access is on pulling it from external should also work.",
      "votes": null
    },
    {
      "id": "787544",
      "postDate": "03/26/2020 21:45:25",
      "content": "<p>If you are interested in working with TPUs, get a TPU or a TPU pod on GCP. You will be able to work privately there.</p>\n\n<p>On Kaggle, here is the current situation with TPUs:\n- private datasets will not work for now through the <code>KaggleDatasets().get_gcs_path()</code> function.\n- you can use a private dataset if you load the data to memory and train from there (using <code>tf.datas.Dataset.from_tensor_slices()</code> or <code>model.fit(numpy_array)</code>). The dataset has to fit in memory though. Training on Python generators will not work.\n- private GCS buckets will not work with TPUs. Theoretically, there is a way through a service account, the auth token and <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/gcs/configure_gcs\">configure_gcs from the tfio</a> lib but it is complicated.</p>",
      "rawMarkdown": "If you are interested in working with TPUs, get a TPU or a TPU pod on GCP. You will be able to work privately there.\n\nOn Kaggle, here is the current situation with TPUs:\n- private datasets will not work for now through the `KaggleDatasets().get_gcs_path()` function.\n- you can use a private dataset if you load the data to memory and train from there (using `tf.datas.Dataset.from_tensor_slices()` or `model.fit(numpy_array)`). The dataset has to fit in memory though. Training on Python generators will not work.\n- private GCS buckets will not work with TPUs. Theoretically, there is a way through a service account, the auth token and [configure_gcs from the tfio](https://www.tensorflow.org/io/api_docs/python/tfio/gcs/configure_gcs) lib but it is complicated.",
      "votes": null
    },
    {
      "id": "787655",
      "postDate": "03/27/2020 00:51:37",
      "content": "<p>But isn't 3 hours too less to other over all data? (I haven't completed digested the TF kernel's but incase of PyTorch, you will get straight OOM if you attempt using all cores with slightly bigger model 😔)!</p>\n\n<p>Seems we should really pick up TF in case we have to use TPUs now and then on Kaggle 🎉🔥</p>\n\n<p>The credentials, you mentioned are service account credentials, right Martin?</p>",
      "rawMarkdown": "But isn't 3 hours too less to other over all data? (I haven't completed digested the TF kernel's but incase of PyTorch, you will get straight OOM if you attempt using all cores with slightly bigger model 😔)!\n\nSeems we should really pick up TF in case we have to use TPUs now and then on Kaggle 🎉🔥\n\nThe credentials, you mentioned are service account credentials, right Martin?",
      "votes": null
    },
    {
      "id": "787715",
      "postDate": "03/27/2020 02:20:30",
      "content": "<p>I'm still trying to get to the bottom of what works what does not in PyTorch-XLA with the team.</p>\n\n<p>For accessing your own private bucket from a TPU on Kaggle, I posted a more complete set of hints <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130326#758575\">here</a>. Not sure it will work.</p>",
      "rawMarkdown": "I'm still trying to get to the bottom of what works what does not in PyTorch-XLA with the team.\n\nFor accessing your own private bucket from a TPU on Kaggle, I posted a more complete set of hints [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130326#758575). Not sure it will work.",
      "votes": null
    },
    {
      "id": "798317",
      "postDate": "04/05/2020 11:55:48",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I have some very weird issues with TPU kernels sometimes. Basically I commit them, and the model is not training at all, staying at same train loss over all epochs. I then go in interactive mode, execute, and everything looks good. I recommit it, and it also looks good. Any idea what that can be?</p>",
      "rawMarkdown": "mgornergoogle I have some very weird issues with TPU kernels sometimes. Basically I commit them, and the model is not training at all, staying at same train loss over all epochs. I then go in interactive mode, execute, and everything looks good. I recommit it, and it also looks good. Any idea what that can be?",
      "votes": null
    },
    {
      "id": "799806",
      "postDate": "04/06/2020 19:11:49",
      "content": "<p>I don't believe there is any difference between commit and interactive mode that could change the behavior of the model. Are you sure this is not a convergence issue in your model. I have seen models before that would only converge occasionally.</p>\n\n<p>If the issue is reproducible, please report again. I will have the team look into it.</p>",
      "rawMarkdown": "I don't believe there is any difference between commit and interactive mode that could change the behavior of the model. Are you sure this is not a convergence issue in your model. I have seen models before that would only converge occasionally.\n\nIf the issue is reproducible, please report again. I will have the team look into it.",
      "votes": null
    },
    {
      "id": "799827",
      "postDate": "04/06/2020 19:57:41",
      "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a>, if I encounter it again I will post it. In the meantime I have another issue that troubles me. My fit is working fine for one epoch, but in second epoch the training gets stuck. I tested it in interactive kernel and it basically does not move after a certain number of batches in second epoch. When I commit the kernel fails after a while and I can see in logs:</p>\n\n<blockquote>\n  <p>5357.6s [NbConvertApp] ERROR | Kernel died while waiting for execute reply.</p>\n</blockquote>\n\n<p>Any idea what that can be? </p>",
      "rawMarkdown": "Thanks @mgornergoogle, if I encounter it again I will post it. In the meantime I have another issue that troubles me. My fit is working fine for one epoch, but in second epoch the training gets stuck. I tested it in interactive kernel and it basically does not move after a certain number of batches in second epoch. When I commit the kernel fails after a while and I can see in logs:\n&gt; 5357.6s [NbConvertApp] ERROR | Kernel died while waiting for execute reply.\n\nAny idea what that can be?",
      "votes": null
    },
    {
      "id": "799949",
      "postDate": "04/06/2020 23:49:14",
      "content": "<p>I have seen some TPU stability issues reported on other threads. They appear to be solved in TF 2.2, to be released very soon. This might be it ?</p>",
      "rawMarkdown": "I have seen some TPU stability issues reported on other threads. They appear to be solved in TF 2.2, to be released very soon. This might be it ?",
      "votes": null
    },
    {
      "id": "800164",
      "postDate": "04/07/2020 06:27:47",
      "content": "<p>Not sure... when will tf 2.2 be available on kaggle kernels?</p>",
      "rawMarkdown": "Not sure... when will tf 2.2 be available on kaggle kernels?",
      "votes": null
    },
    {
      "id": "812763",
      "postDate": "04/19/2020 04:17:33",
      "content": "<p>Does anyone know how to print TPU memory usage from a kernel?</p>",
      "rawMarkdown": "Does anyone know how to print TPU memory usage from a kernel?",
      "votes": null
    },
    {
      "id": "812765",
      "postDate": "04/19/2020 04:18:25",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> saw this in another thread:\n!pip install tensorflow==2.2-rc1</p>",
      "rawMarkdown": "philippsinger saw this in another thread:\n!pip install tensorflow==2.2-rc1",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 785280,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/25/2020 00:05:30",
      "content": "<p>Internet is allowed because TPUs read data from GCS (Google Cloud Storage) so your notebook needs access to that.\nExternal data is allowed in the competition but must respect the external data rules:</p>\n\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 785643,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/25/2020 08:22:07",
          "content": "<p>So using an external API to translate test data is OK? I posted it to the thread.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 787384,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "03/26/2020 18:49:30",
      "content": "<p>Is there some way to use private data in TPU kernels? </p>",
      "votes": null,
      "replies": [
        {
          "id": 787390,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/26/2020 18:52:46",
          "content": "<p>Interesting Question, I think it's possible because if we look at the Getting Started TF notebook, there's a line which calls KaggleDatasets(). something, that should automatically put your data to the nearest GCP bucket available (at least that's what the impression I ha e Collected from Martin's comments somewhere)!\nAlso you can simply use external datasets, no Psi?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787481,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/26/2020 20:18:48",
          "content": "<p>I am debating whether it makes sense to work with TPUs in general because at some point I might need private data I want to import. Maybe encrypting the content and then decrypting in kernel?</p>\n\n<p>But you are right, as Internet access is on pulling it from external should also work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787544,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/26/2020 21:45:25",
          "content": "<p>If you are interested in working with TPUs, get a TPU or a TPU pod on GCP. You will be able to work privately there.</p>\n\n<p>On Kaggle, here is the current situation with TPUs:\n- private datasets will not work for now through the <code>KaggleDatasets().get_gcs_path()</code> function.\n- you can use a private dataset if you load the data to memory and train from there (using <code>tf.datas.Dataset.from_tensor_slices()</code> or <code>model.fit(numpy_array)</code>). The dataset has to fit in memory though. Training on Python generators will not work.\n- private GCS buckets will not work with TPUs. Theoretically, there is a way through a service account, the auth token and <a href=\"https://www.tensorflow.org/io/api_docs/python/tfio/gcs/configure_gcs\">configure_gcs from the tfio</a> lib but it is complicated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787655,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "03/27/2020 00:51:37",
          "content": "<p>But isn't 3 hours too less to other over all data? (I haven't completed digested the TF kernel's but incase of PyTorch, you will get straight OOM if you attempt using all cores with slightly bigger model 😔)!</p>\n\n<p>Seems we should really pick up TF in case we have to use TPUs now and then on Kaggle 🎉🔥</p>\n\n<p>The credentials, you mentioned are service account credentials, right Martin?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 787715,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/27/2020 02:20:30",
          "content": "<p>I'm still trying to get to the bottom of what works what does not in PyTorch-XLA with the team.</p>\n\n<p>For accessing your own private bucket from a TPU on Kaggle, I posted a more complete set of hints <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130326#758575\">here</a>. Not sure it will work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 798317,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/05/2020 11:55:48",
      "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> I have some very weird issues with TPU kernels sometimes. Basically I commit them, and the model is not training at all, staying at same train loss over all epochs. I then go in interactive mode, execute, and everything looks good. I recommit it, and it also looks good. Any idea what that can be?</p>",
      "votes": null,
      "replies": [
        {
          "id": 799806,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/06/2020 19:11:49",
          "content": "<p>I don't believe there is any difference between commit and interactive mode that could change the behavior of the model. Are you sure this is not a convergence issue in your model. I have seen models before that would only converge occasionally.</p>\n\n<p>If the issue is reproducible, please report again. I will have the team look into it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799827,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/06/2020 19:57:41",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a>, if I encounter it again I will post it. In the meantime I have another issue that troubles me. My fit is working fine for one epoch, but in second epoch the training gets stuck. I tested it in interactive kernel and it basically does not move after a certain number of batches in second epoch. When I commit the kernel fails after a while and I can see in logs:</p>\n\n<blockquote>\n  <p>5357.6s [NbConvertApp] ERROR | Kernel died while waiting for execute reply.</p>\n</blockquote>\n\n<p>Any idea what that can be? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 799949,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "04/06/2020 23:49:14",
          "content": "<p>I have seen some TPU stability issues reported on other threads. They appear to be solved in TF 2.2, to be released very soon. This might be it ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800164,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/07/2020 06:27:47",
          "content": "<p>Not sure... when will tf 2.2 be available on kaggle kernels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 812765,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/19/2020 04:18:25",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> saw this in another thread:\n!pip install tensorflow==2.2-rc1</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 812763,
      "author_name": "luohongchen1993",
      "author_url": "",
      "post_date": "04/19/2020 04:17:33",
      "content": "<p>Does anyone know how to print TPU memory usage from a kernel?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "784575": "I think it makes sense to have a dedicated thread for some minor TPU related questions.\n\nI will start:\n\nAs far as I can see, Internet access is always possible in TPU kernels. Isn't that super risky for submission kernels? I can just pull information from anywhere while subbing. Is this allowed?",
    "785280": "Internet is allowed because TPUs read data from GCS (Google Cloud Storage) so your notebook needs access to that.\nExternal data is allowed in the competition but must respect the external data rules:\n\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.",
    "785643": "So using an external API to translate test data is OK? I posted it to the thread.",
    "787384": "Is there some way to use private data in TPU kernels?",
    "787390": "Interesting Question, I think it's possible because if we look at the Getting Started TF notebook, there's a line which calls KaggleDatasets(). something, that should automatically put your data to the nearest GCP bucket available (at least that's what the impression I ha e Collected from Martin's comments somewhere)!\nAlso you can simply use external datasets, no Psi?",
    "787481": "I am debating whether it makes sense to work with TPUs in general because at some point I might need private data I want to import. Maybe encrypting the content and then decrypting in kernel?\n\nBut you are right, as Internet access is on pulling it from external should also work.",
    "787544": "If you are interested in working with TPUs, get a TPU or a TPU pod on GCP. You will be able to work privately there.\n\nOn Kaggle, here is the current situation with TPUs:\n- private datasets will not work for now through the `KaggleDatasets().get_gcs_path()` function.\n- you can use a private dataset if you load the data to memory and train from there (using `tf.datas.Dataset.from_tensor_slices()` or `model.fit(numpy_array)`). The dataset has to fit in memory though. Training on Python generators will not work.\n- private GCS buckets will not work with TPUs. Theoretically, there is a way through a service account, the auth token and [configure_gcs from the tfio](https://www.tensorflow.org/io/api_docs/python/tfio/gcs/configure_gcs) lib but it is complicated.",
    "787655": "But isn't 3 hours too less to other over all data? (I haven't completed digested the TF kernel's but incase of PyTorch, you will get straight OOM if you attempt using all cores with slightly bigger model 😔)!\n\nSeems we should really pick up TF in case we have to use TPUs now and then on Kaggle 🎉🔥\n\nThe credentials, you mentioned are service account credentials, right Martin?",
    "787715": "I'm still trying to get to the bottom of what works what does not in PyTorch-XLA with the team.\n\nFor accessing your own private bucket from a TPU on Kaggle, I posted a more complete set of hints [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130326#758575). Not sure it will work.",
    "798317": "mgornergoogle I have some very weird issues with TPU kernels sometimes. Basically I commit them, and the model is not training at all, staying at same train loss over all epochs. I then go in interactive mode, execute, and everything looks good. I recommit it, and it also looks good. Any idea what that can be?",
    "799806": "I don't believe there is any difference between commit and interactive mode that could change the behavior of the model. Are you sure this is not a convergence issue in your model. I have seen models before that would only converge occasionally.\n\nIf the issue is reproducible, please report again. I will have the team look into it.",
    "799827": "Thanks @mgornergoogle, if I encounter it again I will post it. In the meantime I have another issue that troubles me. My fit is working fine for one epoch, but in second epoch the training gets stuck. I tested it in interactive kernel and it basically does not move after a certain number of batches in second epoch. When I commit the kernel fails after a while and I can see in logs:\n&gt; 5357.6s [NbConvertApp] ERROR | Kernel died while waiting for execute reply.\n\nAny idea what that can be?",
    "799949": "I have seen some TPU stability issues reported on other threads. They appear to be solved in TF 2.2, to be released very soon. This might be it ?",
    "800164": "Not sure... when will tf 2.2 be available on kaggle kernels?",
    "812763": "Does anyone know how to print TPU memory usage from a kernel?",
    "812765": "philippsinger saw this in another thread:\n!pip install tensorflow==2.2-rc1"
  },
  "source": "meta"
}