{
  "id": 103053,
  "title": "TPU code sharing thread",
  "url": "/competitions/recursion-cellular-image-classification/discussion/103053",
  "author_name": "",
  "post_date": "2019-08-06T20:13:04.757258300Z",
  "votes": 12,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Since TPU code can't be shared on Kaggle as easily as kernels, we'd like to provide a centralized thread for swapping public TPU code. In case anyone missed it, here's Recursion's starter code again to kick things off: <a href=\"https://colab.research.google.com/github/recursionpharma/rxrx1-utils/blob/master/notebooks/training.ipynb\">https://colab.research.google.com/github/recursionpharma/rxrx1-utils/blob/master/notebooks/training.ipynb</a></p>\n\n<p>We're looking forward to seeing what you're all up to with your quota time!</p>",
  "messages": [
    {
      "id": "593568",
      "postDate": "08/06/2019 20:13:04",
      "content": "<p>Since TPU code can't be shared on Kaggle as easily as kernels, we'd like to provide a centralized thread for swapping public TPU code. In case anyone missed it, here's Recursion's starter code again to kick things off: <a href=\"https://colab.research.google.com/github/recursionpharma/rxrx1-utils/blob/master/notebooks/training.ipynb\">https://colab.research.google.com/github/recursionpharma/rxrx1-utils/blob/master/notebooks/training.ipynb</a></p>\n\n<p>We're looking forward to seeing what you're all up to with your quota time!</p>",
      "rawMarkdown": "Since TPU code can't be shared on Kaggle as easily as kernels, we'd like to provide a centralized thread for swapping public TPU code. In case anyone missed it, here's Recursion's starter code again to kick things off: https://colab.research.google.com/github/recursionpharma/rxrx1-utils/blob/master/notebooks/training.ipynb\n\nWe're looking forward to seeing what you're all up to with your quota time!",
      "votes": null
    },
    {
      "id": "606275",
      "postDate": "08/23/2019 11:38:59",
      "content": "<p>There is a <a href=\"https://medium.com/@zaharchikishev/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4?source=friends_link&amp;sk=695f9605874c7306837645ff814e7fc9\">short article</a> of my experience running TPU with pytorch so far. But I am less inclined to share code at the moment.</p>",
      "rawMarkdown": "There is a [short article](https://medium.com/@zaharchikishev/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4?source=friends_link&amp;sk=695f9605874c7306837645ff814e7fc9) of my experience running TPU with pytorch so far. But I am less inclined to share code at the moment.",
      "votes": null
    },
    {
      "id": "608290",
      "postDate": "08/26/2019 15:41:48",
      "content": "<p>Thanks! If you're willing to share code and/or additional details after the competition has closed, we'd be very interested in further reading.</p>",
      "rawMarkdown": "Thanks! If you're willing to share code and/or additional details after the competition has closed, we'd be very interested in further reading.",
      "votes": null
    },
    {
      "id": "619595",
      "postDate": "09/06/2019 10:46:08",
      "content": "<p>can anyone tell me if the code necessary for inference on tpus is included in the rxrx.ai repo?</p>",
      "rawMarkdown": "can anyone tell me if the code necessary for inference on tpus is included in the rxrx.ai repo?",
      "votes": null
    },
    {
      "id": "620055",
      "postDate": "09/07/2019 00:22:21",
      "content": "<p>A starter for training has been made available, but not inference.</p>",
      "rawMarkdown": "A starter for training has been made available, but not inference.",
      "votes": null
    },
    {
      "id": "620085",
      "postDate": "09/07/2019 02:22:50",
      "content": "<p>Indeed, the training code for tpus provided by recursion is very helpful. I suppose it is an exercise left to the reader then to figure out how to use the trained model to obtain predictions. I will keep trying then. Thanks.</p>",
      "rawMarkdown": "Indeed, the training code for tpus provided by recursion is very helpful. I suppose it is an exercise left to the reader then to figure out how to use the trained model to obtain predictions. I will keep trying then. Thanks.",
      "votes": null
    },
    {
      "id": "620233",
      "postDate": "09/07/2019 07:20:24",
      "content": "<p>I used the rxrx.ai repo training code (with almost no changes), and then added a <code>resnet_classifier.evaluate()</code> function and a <code>resnet_classifier.predict()</code> function for inference. These, together with <a href=\"https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak\">this trick</a> from the Double Strand team have got me to a LB score of 70% (single model, image recognition only, no ensembling and training from scratch with no pretrained weights). The TPU training is very fast. It takes just over a minute to process a whole epoch of [512, 512, 6] training data through Resnet50 on a google cloud v3-8.</p>\n\n<p>Now I hope to improve that score with a pretrained Resnet, statistical features and ensembling.</p>",
      "rawMarkdown": "I used the rxrx.ai repo training code (with almost no changes), and then added a `resnet_classifier.evaluate()` function and a `resnet_classifier.predict()` function for inference. These, together with [this trick](https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak) from the Double Strand team have got me to a LB score of 70% (single model, image recognition only, no ensembling and training from scratch with no pretrained weights). The TPU training is very fast. It takes just over a minute to process a whole epoch of [512, 512, 6] training data through Resnet50 on a google cloud v3-8.\n\nNow I hope to improve that score with a pretrained Resnet, statistical features and ensembling.",
      "votes": null
    },
    {
      "id": "622747",
      "postDate": "09/10/2019 03:27:18",
      "content": "<p><a href=\"/kenkrige\">@kenkrige</a> Can I ask how many epochs for training?</p>",
      "rawMarkdown": "kenkrige Can I ask how many epochs for training?",
      "votes": null
    },
    {
      "id": "622773",
      "postDate": "09/10/2019 04:42:25",
      "content": "<p><a href=\"/super13579\">@super13579</a> my scores kept improving for about 400 epochs but that was training from random weights. I have not yet used a pretrained model, so I'm sure that would speed it up.</p>",
      "rawMarkdown": "super13579 my scores kept improving for about 400 epochs but that was training from random weights. I have not yet used a pretrained model, so I'm sure that would speed it up.",
      "votes": null
    },
    {
      "id": "629826",
      "postDate": "09/19/2019 08:31:23",
      "content": "<p>HI <a href=\"/sohier\">@sohier</a> , <a href=\"/juliaelliott\">@juliaelliott</a> , <a href=\"/kenkrige\">@kenkrige</a> \nI tried to run the starter code in colab TPU and getting following error - \n<code>\nUnimplementedError: From /job:worker/replica:0/task:0:\nFile system scheme '[local]' not implemented (file: '/gdrive/My Drive/Colab Notebooks/models/model.ckpt-0_temp_3db6d480db1640448b0704187584f502')\n     [[node save/SaveV2 (defined at /content/rxrx1-utils/rxrx/main.py:355) ]]\n</code>\nAny idea why it is not able to save the checkpoint?</p>",
      "rawMarkdown": "HI @sohier , @juliaelliott , @kenkrige \nI tried to run the starter code in colab TPU and getting following error - \n```\nUnimplementedError: From /job:worker/replica:0/task:0:\nFile system scheme '[local]' not implemented (file: '/gdrive/My Drive/Colab Notebooks/models/model.ckpt-0_temp_3db6d480db1640448b0704187584f502')\n\t [[node save/SaveV2 (defined at /content/rxrx1-utils/rxrx/main.py:355) ]]\n```\nAny idea why it is not able to save the checkpoint?",
      "votes": null
    },
    {
      "id": "629834",
      "postDate": "09/19/2019 08:57:28",
      "content": "<p>Hi <a href=\"/kranthi9\">@kranthi9</a> Google Cloud TPUs can only save to a google storage bucket. You need a google cloud platform account and a storage bucket. Then your filename will start with <code>`gs://</code>. That is why you are getting the error <code>File system scheme '[local]' not implemented</code></p>",
      "rawMarkdown": "Hi @kranthi9 Google Cloud TPUs can only save to a google storage bucket. You need a google cloud platform account and a storage bucket. Then your filename will start with ``gs://`. That is why you are getting the error `File system scheme '[local]' not implemented`",
      "votes": null
    },
    {
      "id": "629865",
      "postDate": "09/19/2019 10:56:11",
      "content": "<p>Thanks <a href=\"/kenkrige\">@kenkrige</a> for your quick and informative reply... \nis there any way to run it on Google Colab free TPU instead of Gogle Cloud TPU?</p>",
      "rawMarkdown": "Thanks @kenkrige for your quick and informative reply... \nis there any way to run it on Google Colab free TPU instead of Gogle Cloud TPU?",
      "votes": null
    },
    {
      "id": "629868",
      "postDate": "09/19/2019 11:17:32",
      "content": "<p>You can use the colab TPU but the trained model has to be saved on google cloud storage. What I did was to open a google cloud account and use the $300 they give when you sign up. So far, I have managed to do everything on free credits. A more difficult problem you will encounter is that the code supplied by Recursion for colab only has training code, not prediction code, so I had to add functions for prediction and evaluation.</p>",
      "rawMarkdown": "You can use the colab TPU but the trained model has to be saved on google cloud storage. What I did was to open a google cloud account and use the $300 they give when you sign up. So far, I have managed to do everything on free credits. A more difficult problem you will encounter is that the code supplied by Recursion for colab only has training code, not prediction code, so I had to add functions for prediction and evaluation.",
      "votes": null
    },
    {
      "id": "629871",
      "postDate": "09/19/2019 11:24:22",
      "content": "<p>ok got it..\nI missed the $300 credits deadline of July 22, 2019.. Anyway, Thanks for your explanation <a href=\"/kenkrige\">@kenkrige</a> ... 👍 </p>",
      "rawMarkdown": "ok got it..\nI missed the $300 credits deadline of July 22, 2019.. Anyway, Thanks for your explanation @kenkrige ... 👍",
      "votes": null
    },
    {
      "id": "629873",
      "postDate": "09/19/2019 11:34:14",
      "content": "<p>But you still get $300 credit from google for the first year when you sign up.</p>",
      "rawMarkdown": "But you still get $300 credit from google for the first year when you sign up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 606275,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/23/2019 11:38:59",
      "content": "<p>There is a <a href=\"https://medium.com/@zaharchikishev/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4?source=friends_link&amp;sk=695f9605874c7306837645ff814e7fc9\">short article</a> of my experience running TPU with pytorch so far. But I am less inclined to share code at the moment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 608290,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "08/26/2019 15:41:48",
          "content": "<p>Thanks! If you're willing to share code and/or additional details after the competition has closed, we'd be very interested in further reading.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 619595,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "09/06/2019 10:46:08",
      "content": "<p>can anyone tell me if the code necessary for inference on tpus is included in the rxrx.ai repo?</p>",
      "votes": null,
      "replies": [
        {
          "id": 620055,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "09/07/2019 00:22:21",
          "content": "<p>A starter for training has been made available, but not inference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620085,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "09/07/2019 02:22:50",
          "content": "<p>Indeed, the training code for tpus provided by recursion is very helpful. I suppose it is an exercise left to the reader then to figure out how to use the trained model to obtain predictions. I will keep trying then. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620233,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "09/07/2019 07:20:24",
          "content": "<p>I used the rxrx.ai repo training code (with almost no changes), and then added a <code>resnet_classifier.evaluate()</code> function and a <code>resnet_classifier.predict()</code> function for inference. These, together with <a href=\"https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak\">this trick</a> from the Double Strand team have got me to a LB score of 70% (single model, image recognition only, no ensembling and training from scratch with no pretrained weights). The TPU training is very fast. It takes just over a minute to process a whole epoch of [512, 512, 6] training data through Resnet50 on a google cloud v3-8.</p>\n\n<p>Now I hope to improve that score with a pretrained Resnet, statistical features and ensembling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622747,
          "author_name": "super13579",
          "author_url": "",
          "post_date": "09/10/2019 03:27:18",
          "content": "<p><a href=\"/kenkrige\">@kenkrige</a> Can I ask how many epochs for training?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622773,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "09/10/2019 04:42:25",
          "content": "<p><a href=\"/super13579\">@super13579</a> my scores kept improving for about 400 epochs but that was training from random weights. I have not yet used a pretrained model, so I'm sure that would speed it up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 629826,
      "author_name": "kranthi9",
      "author_url": "",
      "post_date": "09/19/2019 08:31:23",
      "content": "<p>HI <a href=\"/sohier\">@sohier</a> , <a href=\"/juliaelliott\">@juliaelliott</a> , <a href=\"/kenkrige\">@kenkrige</a> \nI tried to run the starter code in colab TPU and getting following error - \n<code>\nUnimplementedError: From /job:worker/replica:0/task:0:\nFile system scheme '[local]' not implemented (file: '/gdrive/My Drive/Colab Notebooks/models/model.ckpt-0_temp_3db6d480db1640448b0704187584f502')\n     [[node save/SaveV2 (defined at /content/rxrx1-utils/rxrx/main.py:355) ]]\n</code>\nAny idea why it is not able to save the checkpoint?</p>",
      "votes": null,
      "replies": [
        {
          "id": 629834,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "09/19/2019 08:57:28",
          "content": "<p>Hi <a href=\"/kranthi9\">@kranthi9</a> Google Cloud TPUs can only save to a google storage bucket. You need a google cloud platform account and a storage bucket. Then your filename will start with <code>`gs://</code>. That is why you are getting the error <code>File system scheme '[local]' not implemented</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 629865,
          "author_name": "kranthi9",
          "author_url": "",
          "post_date": "09/19/2019 10:56:11",
          "content": "<p>Thanks <a href=\"/kenkrige\">@kenkrige</a> for your quick and informative reply... \nis there any way to run it on Google Colab free TPU instead of Gogle Cloud TPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 629868,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "09/19/2019 11:17:32",
          "content": "<p>You can use the colab TPU but the trained model has to be saved on google cloud storage. What I did was to open a google cloud account and use the $300 they give when you sign up. So far, I have managed to do everything on free credits. A more difficult problem you will encounter is that the code supplied by Recursion for colab only has training code, not prediction code, so I had to add functions for prediction and evaluation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 629871,
          "author_name": "kranthi9",
          "author_url": "",
          "post_date": "09/19/2019 11:24:22",
          "content": "<p>ok got it..\nI missed the $300 credits deadline of July 22, 2019.. Anyway, Thanks for your explanation <a href=\"/kenkrige\">@kenkrige</a> ... 👍 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 629873,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "09/19/2019 11:34:14",
          "content": "<p>But you still get $300 credit from google for the first year when you sign up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "593568": "Since TPU code can't be shared on Kaggle as easily as kernels, we'd like to provide a centralized thread for swapping public TPU code. In case anyone missed it, here's Recursion's starter code again to kick things off: https://colab.research.google.com/github/recursionpharma/rxrx1-utils/blob/master/notebooks/training.ipynb\n\nWe're looking forward to seeing what you're all up to with your quota time!",
    "606275": "There is a [short article](https://medium.com/@zaharchikishev/running-pytorch-on-tpu-a-bag-of-tricks-b6d0130bddd4?source=friends_link&amp;sk=695f9605874c7306837645ff814e7fc9) of my experience running TPU with pytorch so far. But I am less inclined to share code at the moment.",
    "608290": "Thanks! If you're willing to share code and/or additional details after the competition has closed, we'd be very interested in further reading.",
    "619595": "can anyone tell me if the code necessary for inference on tpus is included in the rxrx.ai repo?",
    "620055": "A starter for training has been made available, but not inference.",
    "620085": "Indeed, the training code for tpus provided by recursion is very helpful. I suppose it is an exercise left to the reader then to figure out how to use the trained model to obtain predictions. I will keep trying then. Thanks.",
    "620233": "I used the rxrx.ai repo training code (with almost no changes), and then added a `resnet_classifier.evaluate()` function and a `resnet_classifier.predict()` function for inference. These, together with [this trick](https://www.kaggle.com/zaharch/keras-model-boosted-with-plates-leak) from the Double Strand team have got me to a LB score of 70% (single model, image recognition only, no ensembling and training from scratch with no pretrained weights). The TPU training is very fast. It takes just over a minute to process a whole epoch of [512, 512, 6] training data through Resnet50 on a google cloud v3-8.\n\nNow I hope to improve that score with a pretrained Resnet, statistical features and ensembling.",
    "622747": "kenkrige Can I ask how many epochs for training?",
    "622773": "super13579 my scores kept improving for about 400 epochs but that was training from random weights. I have not yet used a pretrained model, so I'm sure that would speed it up.",
    "629826": "HI @sohier , @juliaelliott , @kenkrige \nI tried to run the starter code in colab TPU and getting following error - \n```\nUnimplementedError: From /job:worker/replica:0/task:0:\nFile system scheme '[local]' not implemented (file: '/gdrive/My Drive/Colab Notebooks/models/model.ckpt-0_temp_3db6d480db1640448b0704187584f502')\n\t [[node save/SaveV2 (defined at /content/rxrx1-utils/rxrx/main.py:355) ]]\n```\nAny idea why it is not able to save the checkpoint?",
    "629834": "Hi @kranthi9 Google Cloud TPUs can only save to a google storage bucket. You need a google cloud platform account and a storage bucket. Then your filename will start with ``gs://`. That is why you are getting the error `File system scheme '[local]' not implemented`",
    "629865": "Thanks @kenkrige for your quick and informative reply... \nis there any way to run it on Google Colab free TPU instead of Gogle Cloud TPU?",
    "629868": "You can use the colab TPU but the trained model has to be saved on google cloud storage. What I did was to open a google cloud account and use the $300 they give when you sign up. So far, I have managed to do everything on free credits. A more difficult problem you will encounter is that the code supplied by Recursion for colab only has training code, not prediction code, so I had to add functions for prediction and evaluation.",
    "629871": "ok got it..\nI missed the $300 credits deadline of July 22, 2019.. Anyway, Thanks for your explanation @kenkrige ... 👍",
    "629873": "But you still get $300 credit from google for the first year when you sign up."
  },
  "source": "meta"
}