{
  "id": 134161,
  "title": "Bengali on TPUs [1 model 1 fold -> 0.9708]",
  "url": "/competitions/bengaliai-cv19/discussion/134161",
  "author_name": "See--",
  "post_date": "2020-03-06T11:04:43.352000",
  "votes": 83,
  "comment_count": 43,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n\n<p>Thanks to Kaggle we can now use super fast TPUs for free! I wanted to apply them in this challenge so I created the following short series of Notebooks:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/seesee/1-create-tfrecords\">1. Create TFRecords</a></li>\n</ul>\n\n<p>This is a two-step process. The images are first converted to <code>.png</code> format and then converted to <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\"><code>TFRecords</code></a>(thanks to <a href=\"/mgornergoogle\">@mgornergoogle</a> and <a href=\"/rsmits\">@rsmits</a> for the code). Creating <code>TFRecords</code> is recommended for datasets that don't fit into memory. Note that you can skip this step as I already created the following dataset: <a href=\"https://www.kaggle.com/seesee/bengali-tfrecords-v010\">https://www.kaggle.com/seesee/bengali-tfrecords-v010</a></p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/seesee/2-train\">2. Train</a></li>\n</ul>\n\n<p>Efficientnet-B3 is trained for 25 epochs. You can easily switch the <code>backbone</code> thanks to the <a href=\"https://github.com/qubvel/efficientnet\">efficientnet library</a>. There are 2 parts that are not \"vanilla\": I am using <code>mixup</code> augmentation and I train with <code>softmax</code> and then decode the predictions. There is a ton that can be improved. For example, you can try <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132894\">these weights</a>, a non-constant learning rate schedule, multiple folds, etc. The training takes less than 2 hours and I get 0.9708 on LB.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/seesee/3-submit\">3. Submit</a></li>\n</ul>\n\n<p>TPU Notebooks are not allowed for submission so I created a new Notebook with GPU accelerator. The inference takes ~30 minutes.</p>\n\n<p>I hope this makes it easier to work with TPUs on Kaggle!</p>",
  "messages": [
    {
      "id": 765194,
      "postDate": "2020-03-06T11:04:43.353Z",
      "content": "<p>Hi everyone,</p>\n\n<p>Thanks to Kaggle we can now use super fast TPUs for free! I wanted to apply them in this challenge so I created the following short series of Notebooks:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/seesee/1-create-tfrecords\">1. Create TFRecords</a></li>\n</ul>\n\n<p>This is a two-step process. The images are first converted to <code>.png</code> format and then converted to <a href=\"https://www.tensorflow.org/tutorials/load_data/tfrecord\"><code>TFRecords</code></a>(thanks to <a href=\"/mgornergoogle\">@mgornergoogle</a> and <a href=\"/rsmits\">@rsmits</a> for the code). Creating <code>TFRecords</code> is recommended for datasets that don't fit into memory. Note that you can skip this step as I already created the following dataset: <a href=\"https://www.kaggle.com/seesee/bengali-tfrecords-v010\">https://www.kaggle.com/seesee/bengali-tfrecords-v010</a></p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/seesee/2-train\">2. Train</a></li>\n</ul>\n\n<p>Efficientnet-B3 is trained for 25 epochs. You can easily switch the <code>backbone</code> thanks to the <a href=\"https://github.com/qubvel/efficientnet\">efficientnet library</a>. There are 2 parts that are not \"vanilla\": I am using <code>mixup</code> augmentation and I train with <code>softmax</code> and then decode the predictions. There is a ton that can be improved. For example, you can try <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/132894\">these weights</a>, a non-constant learning rate schedule, multiple folds, etc. The training takes less than 2 hours and I get 0.9708 on LB.</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/seesee/3-submit\">3. Submit</a></li>\n</ul>\n\n<p>TPU Notebooks are not allowed for submission so I created a new Notebook with GPU accelerator. The inference takes ~30 minutes.</p>\n\n<p>I hope this makes it easier to work with TPUs on Kaggle!</p>",
      "rawMarkdown": "Hi everyone,\n\nThanks to Kaggle we can now use super fast TPUs for free! I wanted to apply them in this challenge so I created the following short series of Notebooks:\n\n- [1. Create TFRecords](https://www.kaggle.com/seesee/1-create-tfrecords)\n\nThis is a two-step process. The images are first converted to `.png` format and then converted to [`TFRecords`](https://www.tensorflow.org/tutorials/load_data/tfrecord)(thanks to @mgornergoogle and @rsmits for the code). Creating `TFRecords` is recommended for datasets that don't fit into memory. Note that you can skip this step as I already created the following dataset: https://www.kaggle.com/seesee/bengali-tfrecords-v010\n\n- [2. Train](https://www.kaggle.com/seesee/2-train)\n\nEfficientnet-B3 is trained for 25 epochs. You can easily switch the `backbone` thanks to the [efficientnet library](https://github.com/qubvel/efficientnet). There are 2 parts that are not \"vanilla\": I am using `mixup` augmentation and I train with `softmax` and then decode the predictions. There is a ton that can be improved. For example, you can try [these weights](https://www.kaggle.com/c/bengaliai-cv19/discussion/132894), a non-constant learning rate schedule, multiple folds, etc. The training takes less than 2 hours and I get 0.9708 on LB.\n\n- [3. Submit](https://www.kaggle.com/seesee/3-submit)\n\nTPU Notebooks are not allowed for submission so I created a new Notebook with GPU accelerator. The inference takes ~30 minutes.\n\nI hope this makes it easier to work with TPUs on Kaggle!",
      "votes": 83
    },
    {
      "id": 770736,
      "postDate": "2020-03-13T10:38:03.213Z",
      "content": "<p>Many thanks, I need a support on TPU for my research!</p>",
      "rawMarkdown": "Many thanks, I need a support on TPU for my research!",
      "votes": 1
    },
    {
      "id": 766631,
      "postDate": "2020-03-08T13:36:27.387Z",
      "content": "<p>Have anyone tried to train on colab TPU? I ran out my TPU quora this week, and then I started to work on reproduce this on google colab TPU, but keep failing....\nI still encounter this issue even though it is said fixed. \n<a href=\"https://github.com/googlecolab/colabtools/issues/808\">https://github.com/googlecolab/colabtools/issues/808</a></p>",
      "rawMarkdown": "Have anyone tried to train on colab TPU? I ran out my TPU quora this week, and then I started to work on reproduce this on google colab TPU, but keep failing....\nI still encounter this issue even though it is said fixed. \nhttps://github.com/googlecolab/colabtools/issues/808",
      "votes": 1,
      "replies": [
        {
          "id": 766649,
          "postDate": "2020-03-08T14:13:34.530Z",
          "content": "<p>Did you copy the dataset to your own GCS bucket? I am not sure if you can access Kaggle datasets outside of Kaggle (i.e. Colab). I recall that for my local machine I also had to run:</p>\n\n<p><code>export GOOGLE_APPLICATION_CREDENTIALS=\"/home/MY_USER/.config/gcloud/legacy_credentials/MY_MAIL@XY.com/adc.json\"</code></p>",
          "rawMarkdown": "Did you copy the dataset to your own GCS bucket? I am not sure if you can access Kaggle datasets outside of Kaggle (i.e. Colab). I recall that for my local machine I also had to run:\n\n`export GOOGLE_APPLICATION_CREDENTIALS=\"/home/MY_USER/.config/gcloud/legacy_credentials/MY_MAIL@XY.com/adc.json\"`"
        },
        {
          "id": 766680,
          "postDate": "2020-03-08T14:56:28.380Z",
          "content": "<p><a href=\"/bamps53\">@bamps53</a> , try downloading the datasets via the Kaggle API instead of using your GCS bucket. It should work as I have done been doing that to train my models be it in GPU and I do not think it would be differemt for TPU.</p>\n\n<p><a href=\"/seesee\">@seesee</a> , yes you can download kaggle competition data or any Kaggle dataset in colab by using the kaggle API.</p>",
          "rawMarkdown": "@bamps53 , try downloading the datasets via the Kaggle API instead of using your GCS bucket. It should work as I have done been doing that to train my models be it in GPU and I do not think it would be differemt for TPU.\n\n@seesee , yes you can download kaggle competition data or any Kaggle dataset in colab by using the kaggle API."
        },
        {
          "id": 766681,
          "postDate": "2020-03-08T14:58:15.597Z",
          "content": "<p>i will try today. Regarding data - I plan to upload it to google cloud, where u can use it in colab. </p>",
          "rawMarkdown": "i will try today. Regarding data - I plan to upload it to google cloud, where u can use it in colab. "
        },
        {
          "id": 766683,
          "postDate": "2020-03-08T15:04:51.943Z",
          "content": "<p><a href=\"/serg132003\">@serg132003</a>, I meant you can download the data into your colab notebook and use it there.</p>",
          "rawMarkdown": "@serg132003, I meant you can download the data into your colab notebook and use it there."
        },
        {
          "id": 766687,
          "postDate": "2020-03-08T15:12:00.833Z",
          "content": "<p>It is different for TPUs. You <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#cannot_use_local_filesystem\">can't use the local filesystem</a>. The data has to go to GCS.</p>",
          "rawMarkdown": "It is different for TPUs. You [can't use the local filesystem](https://cloud.google.com/tpu/docs/troubleshooting#cannot_use_local_filesystem). The data has to go to GCS."
        },
        {
          "id": 766690,
          "postDate": "2020-03-08T15:14:17.373Z",
          "content": "<p><a href=\"/seesee\">@seesee</a> \nYes, I did. I also changed the code around data reading like that. \n<code>\nds_path = 'gs://MY_STORAGE_NAME/'\ntrain_fns = tf.io.gfile.glob(os.path.join(ds_path, 'train*.tfrec'))\n</code>\nThanks, I'll try that one.\nThat means you can train model with your tf records in your cloud storage and cloud TPU instance, right??</p>\n\n<p><a href=\"/sheriytm\">@sheriytm</a> \nI think it's different for TPU because we need to put the data on GCS, not local storage.\nI actually downloaded the data with using kaggle API and uploaded to my cloud storage.</p>",
          "rawMarkdown": "@seesee \nYes, I did. I also changed the code around data reading like that. \n`\nds_path = 'gs://MY_STORAGE_NAME/'\ntrain_fns = tf.io.gfile.glob(os.path.join(ds_path, 'train*.tfrec'))\n`\nThanks, I'll try that one.\nThat means you can train model with your tf records in your cloud storage and cloud TPU instance, right??\n\n@sheriytm \nI think it's different for TPU because we need to put the data on GCS, not local storage.\nI actually downloaded the data with using kaggle API and uploaded to my cloud storage."
        },
        {
          "id": 766702,
          "postDate": "2020-03-08T15:26:23.393Z",
          "content": "<p>Thanks @bamps, and <a href=\"/seesee\">@seesee</a>. I did not know that it is  different for TPU.</p>",
          "rawMarkdown": "Thanks @bamps, and @seesee. I did not know that it is  different for TPU."
        },
        {
          "id": 770649,
          "postDate": "2020-03-13T08:11:37.060Z",
          "content": "<p>Hello <a href=\"/bamps53\">@bamps53</a></p>\n\n<p>In case if you are looking for the solution without uploading dataset into your own GC then please try the following steps,</p>\n\n<p>From Kaggle kernel, take your tf record dataset ds_path then you can use it directly into your google colab notebook.</p>\n\n<p><code>from kaggle_datasets import KaggleDatasets</code>\n<code>ds_path = KaggleDatasets().get_gcs_path('bengali-tfrecords-v010')</code></p>\n\n<p>In Colab,</p>\n\n<p><code>ds_path = 'gs://bucket-name/records/' # whatever you get from the kaggle kernel\ntrain_fns = tf.io.gfile.glob(os.path.join(ds_path, 'train*.tfrec'))</code></p>\n\n<p>GC bucket link looks like below,</p>\n\n<p><code>gs://kds-a342e63020315575e027e0fd6e74eba704b178cad83fd7a524e96012</code></p>\n\n<p>Everything else is same, you can train on TPU in colab.</p>\n\n<p>Reference: <a href=\"https://www.kaggle.com/paultimothymooney/how-to-retrieve-gcs-paths-from-kaggle-datasets\">https://www.kaggle.com/paultimothymooney/how-to-retrieve-gcs-paths-from-kaggle-datasets</a></p>\n\n<p>Thank you!</p>",
          "rawMarkdown": "Hello @bamps53\n\nIn case if you are looking for the solution without uploading dataset into your own GC then please try the following steps,\n\nFrom Kaggle kernel, take your tf record dataset ds_path then you can use it directly into your google colab notebook.\n\n`from kaggle_datasets import KaggleDatasets`\n`ds_path = KaggleDatasets().get_gcs_path('bengali-tfrecords-v010')`\n\nIn Colab,\n\n`ds_path = 'gs://bucket-name/records/' # whatever you get from the kaggle kernel\ntrain_fns = tf.io.gfile.glob(os.path.join(ds_path, 'train*.tfrec'))`\n\nGC bucket link looks like below,\n\n`gs://kds-a342e63020315575e027e0fd6e74eba704b178cad83fd7a524e96012`\n\nEverything else is same, you can train on TPU in colab.\n\nReference: https://www.kaggle.com/paultimothymooney/how-to-retrieve-gcs-paths-from-kaggle-datasets\n\nThank you!",
          "votes": 4
        },
        {
          "id": 770685,
          "postDate": "2020-03-13T09:01:37.100Z",
          "content": "<p><a href=\"/anandselvadurai\">@anandselvadurai</a> \nThank you very much. It work for me.</p>",
          "rawMarkdown": "@anandselvadurai \nThank you very much. It work for me.",
          "votes": 1
        },
        {
          "id": 770871,
          "postDate": "2020-03-13T13:58:04.870Z",
          "content": "<p><a href=\"/anandselvadurai\">@anandselvadurai</a> Thanks, that works!!!</p>",
          "rawMarkdown": "@anandselvadurai Thanks, that works!!!",
          "votes": 1
        },
        {
          "id": 771364,
          "postDate": "2020-03-14T04:47:35.800Z",
          "content": "<p>Thanks a lot man!</p>",
          "rawMarkdown": "Thanks a lot man!",
          "votes": 1
        }
      ]
    },
    {
      "id": 766054,
      "postDate": "2020-03-07T15:48:02.657Z",
      "content": "<p>Hi! Would you explain what is \"tuple_map\" in submit, and how did you get it?\nThank you</p>",
      "rawMarkdown": "Hi! Would you explain what is \"tuple_map\" in submit, and how did you get it?\nThank you",
      "votes": 1,
      "replies": [
        {
          "id": 766653,
          "postDate": "2020-03-08T14:17:01.293Z",
          "content": "<p>It maps from (root, vowel, consonant) to a unique int. I used it to just have a single output.</p>",
          "rawMarkdown": "It maps from (root, vowel, consonant) to a unique int. I used it to just have a single output.",
          "votes": 3
        },
        {
          "id": 767566,
          "postDate": "2020-03-09T20:21:50.387Z",
          "content": "<p>Was about to ask this</p>",
          "rawMarkdown": "Was about to ask this"
        }
      ]
    },
    {
      "id": 765771,
      "postDate": "2020-03-07T04:21:24.433Z",
      "content": "<p><a href=\"/seesee\">@seesee</a>  Thanks for sharing this is really helpful!</p>\n\n<p>When I tried my idea based on your kernel, sometimes the kernel was suddenly killed and there was no error message.\nI ran the code with GPU and it ran successfully without any problem.\nI experienced almost same situation when I used pytorch xla.</p>\n\n<p>So, my question is, is there any better way to debug code with TPU??</p>\n\n<p>Thanks in advance!</p>",
      "rawMarkdown": "@seesee  Thanks for sharing this is really helpful!\n\nWhen I tried my idea based on your kernel, sometimes the kernel was suddenly killed and there was no error message.\nI ran the code with GPU and it ran successfully without any problem.\nI experienced almost same situation when I used pytorch xla.\n\nSo, my question is, is there any better way to debug code with TPU??\n\nThanks in advance!",
      "votes": 1,
      "replies": [
        {
          "id": 765861,
          "postDate": "2020-03-07T08:32:18.467Z",
          "content": "<p>I always got at least an error message. Things were quite stable so far. Do you use the interactive mode for debugging? Can you share what you changed?</p>",
          "rawMarkdown": "I always got at least an error message. Things were quite stable so far. Do you use the interactive mode for debugging? Can you share what you changed?"
        },
        {
          "id": 766628,
          "postDate": "2020-03-08T13:33:58.153Z",
          "content": "<p>Thanks, finally I found my model graph was disconnected somewhere. But it seems still strange for me that sometimes it tells me it's disconnected, but sometimes just shutdown the kernel...\nAnyway, now I need to deal with 'Submission CSV Not Found' and 'Submission Scoring Error' in submit kernel:(</p>",
          "rawMarkdown": "Thanks, finally I found my model graph was disconnected somewhere. But it seems still strange for me that sometimes it tells me it's disconnected, but sometimes just shutdown the kernel...\nAnyway, now I need to deal with 'Submission CSV Not Found' and 'Submission Scoring Error' in submit kernel:(\n"
        }
      ]
    },
    {
      "id": 765706,
      "postDate": "2020-03-07T01:37:50.120Z",
      "content": "<p>This is amazing! Thanks for the great work.</p>",
      "rawMarkdown": "This is amazing! Thanks for the great work.",
      "votes": 1
    },
    {
      "id": 765434,
      "postDate": "2020-03-06T16:08:54.440Z",
      "content": "<p><a href=\"/seesee\">@seesee</a>, thanks so much for sharing. I hardly use my TPU quota and you just gave me a reason to experiment.</p>",
      "rawMarkdown": "@seesee, thanks so much for sharing. I hardly use my TPU quota and you just gave me a reason to experiment.",
      "votes": 1
    },
    {
      "id": 765219,
      "postDate": "2020-03-06T11:52:28.050Z",
      "content": "<p>Love you！brother goose🤓 </p>",
      "rawMarkdown": "Love you！brother goose🤓 ",
      "votes": 1
    },
    {
      "id": 765215,
      "postDate": "2020-03-06T11:43:18.940Z",
      "content": "<p>thx for the contribution. I wonder what's your motivation to use <code>one head + softmax</code> here. Any idea how it's going to perform on unseen combinations?</p>",
      "rawMarkdown": "thx for the contribution. I wonder what's your motivation to use `one head + softmax` here. Any idea how it's going to perform on unseen combinations?",
      "votes": 1,
      "replies": [
        {
          "id": 765311,
          "postDate": "2020-03-06T13:31:34.940Z",
          "content": "<p>&gt;  Any idea how it's going to perform on unseen combinations?</p>\n\n<p>Probably not that well :). If you just decode the predictions using <code>argmax</code> the model is not able to predict unseen combinations.</p>\n\n<p>Edit: The motivation was to create a minimal working approach. There is a lot that can be improved.</p>",
          "rawMarkdown": "&gt;  Any idea how it's going to perform on unseen combinations?\n\nProbably not that well :). If you just decode the predictions using `argmax` the model is not able to predict unseen combinations.\n\nEdit: The motivation was to create a minimal working approach. There is a lot that can be improved.",
          "votes": 2
        },
        {
          "id": 767569,
          "postDate": "2020-03-09T20:23:12.043Z",
          "content": "<p>Can you mention some new idea other than changing architecture to ENB7?</p>",
          "rawMarkdown": "Can you mention some new idea other than changing architecture to ENB7?",
          "votes": 1
        }
      ]
    },
    {
      "id": 765666,
      "postDate": "2020-03-07T00:04:00.457Z",
      "content": "<p>just Wow many thanks <a href=\"/seesee\">@seesee</a> , keep the great work</p>",
      "rawMarkdown": "just Wow many thanks @seesee , keep the great work",
      "votes": 2
    },
    {
      "id": 770146,
      "postDate": "2020-03-12T16:02:36.643Z",
      "content": "<p>Good day! When augmentation should be applied - to the original data, or somehow after they are TFR conversion?</p>",
      "rawMarkdown": "Good day! When augmentation should be applied - to the original data, or somehow after they are TFR conversion?",
      "replies": [
        {
          "id": 770743,
          "postDate": "2020-03-13T10:44:05.833Z",
          "content": "<p>The augmentation should be applied after the conversion to records. You can check my <a href=\"https://www.kaggle.com/seesee/2-train\">training Notebook</a> for an example how it can be done: <code>train_ds = train_ds.map(lambda a, b: mixup(a, b, args.batch_size), num_parallel_calls=AUTO)</code>. This augmentation takes a whole batch as input. To augment single images you'd move the augmentation before the <code>ds.batch</code> step.</p>",
          "rawMarkdown": "The augmentation should be applied after the conversion to records. You can check my [training Notebook](https://www.kaggle.com/seesee/2-train) for an example how it can be done: `train_ds = train_ds.map(lambda a, b: mixup(a, b, args.batch_size), num_parallel_calls=AUTO)`. This augmentation takes a whole batch as input. To augment single images you'd move the augmentation before the `ds.batch` step.\n"
        }
      ]
    },
    {
      "id": 767756,
      "postDate": "2020-03-10T04:07:09.283Z",
      "content": "<p>Don't do inference with CPU for 3rd notebooks</p>",
      "rawMarkdown": "Don't do inference with CPU for 3rd notebooks"
    },
    {
      "id": 766691,
      "postDate": "2020-03-08T15:14:37.803Z",
      "content": "<p>Very Nice!</p>",
      "rawMarkdown": "Very Nice!"
    },
    {
      "id": 766346,
      "postDate": "2020-03-08T03:22:49.003Z",
      "content": "<p>Thank you for sharing! We can dramatically reduce training time. </p>",
      "rawMarkdown": "Thank you for sharing! We can dramatically reduce training time. "
    },
    {
      "id": 766076,
      "postDate": "2020-03-07T16:20:45.627Z",
      "content": "<p>its nice</p>",
      "rawMarkdown": "its nice\n"
    },
    {
      "id": 765734,
      "postDate": "2020-03-07T02:16:59.537Z",
      "content": "<p>I was thinking to try on TPU in this competition but couldn't get enough time. I have tried CapsNet though, but no luck.  Anyway, thanks for sharing, amazing work indeed. </p>",
      "rawMarkdown": "I was thinking to try on TPU in this competition but couldn't get enough time. I have tried CapsNet though, but no luck.  Anyway, thanks for sharing, amazing work indeed. "
    },
    {
      "id": 772154,
      "postDate": "2020-03-15T05:20:12.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 770262,
      "postDate": "2020-03-12T18:42:23.823Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 765471,
      "postDate": "2020-03-06T16:55:52.523Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 765315,
      "postDate": "2020-03-06T13:34:38.600Z",
      "content": "<p>Many thanks! :-)</p>",
      "rawMarkdown": "Many thanks! :-)",
      "votes": 1
    },
    {
      "id": 765225,
      "postDate": "2020-03-06T11:55:25.900Z",
      "content": "<p>Great stuff as usual! Thank you!</p>",
      "rawMarkdown": "Great stuff as usual! Thank you!",
      "votes": 2
    },
    {
      "id": 770077,
      "postDate": "2020-03-12T14:43:47.153Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 769979,
      "postDate": "2020-03-12T13:05:32.353Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 768647,
      "postDate": "2020-03-11T03:57:17.987Z",
      "content": "<p>Thanks for sharing. Will try TPU.</p>",
      "rawMarkdown": "Thanks for sharing. Will try TPU."
    },
    {
      "id": 768051,
      "postDate": "2020-03-10T12:02:18.583Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 767856,
      "postDate": "2020-03-10T07:09:15.003Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 770736,
      "author_name": "Salman Chen",
      "author_url": "",
      "post_date": "2020-03-13T10:38:03.213000",
      "content": "<p>Many thanks, I need a support on TPU for my research!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 766631,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-03-08T13:36:27.387000",
      "content": "<p>Have anyone tried to train on colab TPU? I ran out my TPU quora this week, and then I started to work on reproduce this on google colab TPU, but keep failing....\nI still encounter this issue even though it is said fixed. \n<a href=\"https://github.com/googlecolab/colabtools/issues/808\">https://github.com/googlecolab/colabtools/issues/808</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 766649,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-03-08T14:13:34.530000",
          "content": "<p>Did you copy the dataset to your own GCS bucket? I am not sure if you can access Kaggle datasets outside of Kaggle (i.e. Colab). I recall that for my local machine I also had to run:</p>\n\n<p><code>export GOOGLE_APPLICATION_CREDENTIALS=\"/home/MY_USER/.config/gcloud/legacy_credentials/MY_MAIL@XY.com/adc.json\"</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766680,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-03-08T14:56:28.380000",
          "content": "<p><a href=\"/bamps53\">@bamps53</a> , try downloading the datasets via the Kaggle API instead of using your GCS bucket. It should work as I have done been doing that to train my models be it in GPU and I do not think it would be differemt for TPU.</p>\n\n<p><a href=\"/seesee\">@seesee</a> , yes you can download kaggle competition data or any Kaggle dataset in colab by using the kaggle API.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766681,
          "author_name": "Serge",
          "author_url": "",
          "post_date": "2020-03-08T14:58:15.597000",
          "content": "<p>i will try today. Regarding data - I plan to upload it to google cloud, where u can use it in colab. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766683,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-03-08T15:04:51.943000",
          "content": "<p><a href=\"/serg132003\">@serg132003</a>, I meant you can download the data into your colab notebook and use it there.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766687,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-03-08T15:12:00.833000",
          "content": "<p>It is different for TPUs. You <a href=\"https://cloud.google.com/tpu/docs/troubleshooting#cannot_use_local_filesystem\">can't use the local filesystem</a>. The data has to go to GCS.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766690,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-08T15:14:17.373000",
          "content": "<p><a href=\"/seesee\">@seesee</a> \nYes, I did. I also changed the code around data reading like that. \n<code>\nds_path = 'gs://MY_STORAGE_NAME/'\ntrain_fns = tf.io.gfile.glob(os.path.join(ds_path, 'train*.tfrec'))\n</code>\nThanks, I'll try that one.\nThat means you can train model with your tf records in your cloud storage and cloud TPU instance, right??</p>\n\n<p><a href=\"/sheriytm\">@sheriytm</a> \nI think it's different for TPU because we need to put the data on GCS, not local storage.\nI actually downloaded the data with using kaggle API and uploaded to my cloud storage.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766702,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2020-03-08T15:26:23.393000",
          "content": "<p>Thanks @bamps, and <a href=\"/seesee\">@seesee</a>. I did not know that it is  different for TPU.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770649,
          "author_name": "Anand Selvadurai",
          "author_url": "",
          "post_date": "2020-03-13T08:11:37.060000",
          "content": "<p>Hello <a href=\"/bamps53\">@bamps53</a></p>\n\n<p>In case if you are looking for the solution without uploading dataset into your own GC then please try the following steps,</p>\n\n<p>From Kaggle kernel, take your tf record dataset ds_path then you can use it directly into your google colab notebook.</p>\n\n<p><code>from kaggle_datasets import KaggleDatasets</code>\n<code>ds_path = KaggleDatasets().get_gcs_path('bengali-tfrecords-v010')</code></p>\n\n<p>In Colab,</p>\n\n<p><code>ds_path = 'gs://bucket-name/records/' # whatever you get from the kaggle kernel\ntrain_fns = tf.io.gfile.glob(os.path.join(ds_path, 'train*.tfrec'))</code></p>\n\n<p>GC bucket link looks like below,</p>\n\n<p><code>gs://kds-a342e63020315575e027e0fd6e74eba704b178cad83fd7a524e96012</code></p>\n\n<p>Everything else is same, you can train on TPU in colab.</p>\n\n<p>Reference: <a href=\"https://www.kaggle.com/paultimothymooney/how-to-retrieve-gcs-paths-from-kaggle-datasets\">https://www.kaggle.com/paultimothymooney/how-to-retrieve-gcs-paths-from-kaggle-datasets</a></p>\n\n<p>Thank you!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 770685,
          "author_name": "susuky",
          "author_url": "",
          "post_date": "2020-03-13T09:01:37.100000",
          "content": "<p><a href=\"/anandselvadurai\">@anandselvadurai</a> \nThank you very much. It work for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 770871,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-13T13:58:04.870000",
          "content": "<p><a href=\"/anandselvadurai\">@anandselvadurai</a> Thanks, that works!!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 771364,
          "author_name": "sajwankit",
          "author_url": "",
          "post_date": "2020-03-14T04:47:35.800000",
          "content": "<p>Thanks a lot man!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 766054,
      "author_name": "Serge",
      "author_url": "",
      "post_date": "2020-03-07T15:48:02.657000",
      "content": "<p>Hi! Would you explain what is \"tuple_map\" in submit, and how did you get it?\nThank you</p>",
      "votes": 1,
      "replies": [
        {
          "id": 766653,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-03-08T14:17:01.293000",
          "content": "<p>It maps from (root, vowel, consonant) to a unique int. I used it to just have a single output.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 767566,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-03-09T20:21:50.387000",
          "content": "<p>Was about to ask this</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 765771,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-03-07T04:21:24.433000",
      "content": "<p><a href=\"/seesee\">@seesee</a>  Thanks for sharing this is really helpful!</p>\n\n<p>When I tried my idea based on your kernel, sometimes the kernel was suddenly killed and there was no error message.\nI ran the code with GPU and it ran successfully without any problem.\nI experienced almost same situation when I used pytorch xla.</p>\n\n<p>So, my question is, is there any better way to debug code with TPU??</p>\n\n<p>Thanks in advance!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 765861,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-03-07T08:32:18.467000",
          "content": "<p>I always got at least an error message. Things were quite stable so far. Do you use the interactive mode for debugging? Can you share what you changed?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766628,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-03-08T13:33:58.153000",
          "content": "<p>Thanks, finally I found my model graph was disconnected somewhere. But it seems still strange for me that sometimes it tells me it's disconnected, but sometimes just shutdown the kernel...\nAnyway, now I need to deal with 'Submission CSV Not Found' and 'Submission Scoring Error' in submit kernel:(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 765706,
      "author_name": "Tsai29",
      "author_url": "",
      "post_date": "2020-03-07T01:37:50.120000",
      "content": "<p>This is amazing! Thanks for the great work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 765434,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-03-06T16:08:54.440000",
      "content": "<p><a href=\"/seesee\">@seesee</a>, thanks so much for sharing. I hardly use my TPU quota and you just gave me a reason to experiment.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 765219,
      "author_name": "HZD",
      "author_url": "",
      "post_date": "2020-03-06T11:52:28.050000",
      "content": "<p>Love you！brother goose🤓 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 765215,
      "author_name": "Miroslav Valan",
      "author_url": "",
      "post_date": "2020-03-06T11:43:18.940000",
      "content": "<p>thx for the contribution. I wonder what's your motivation to use <code>one head + softmax</code> here. Any idea how it's going to perform on unseen combinations?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 765311,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-03-06T13:31:34.940000",
          "content": "<p>&gt;  Any idea how it's going to perform on unseen combinations?</p>\n\n<p>Probably not that well :). If you just decode the predictions using <code>argmax</code> the model is not able to predict unseen combinations.</p>\n\n<p>Edit: The motivation was to create a minimal working approach. There is a lot that can be improved.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 767569,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-03-09T20:23:12.043000",
          "content": "<p>Can you mention some new idea other than changing architecture to ENB7?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 765666,
      "author_name": "Yassine Hamdaoui",
      "author_url": "",
      "post_date": "2020-03-07T00:04:00.457000",
      "content": "<p>just Wow many thanks <a href=\"/seesee\">@seesee</a> , keep the great work</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 770146,
      "author_name": "Serge",
      "author_url": "",
      "post_date": "2020-03-12T16:02:36.643000",
      "content": "<p>Good day! When augmentation should be applied - to the original data, or somehow after they are TFR conversion?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 770743,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-03-13T10:44:05.833000",
          "content": "<p>The augmentation should be applied after the conversion to records. You can check my <a href=\"https://www.kaggle.com/seesee/2-train\">training Notebook</a> for an example how it can be done: <code>train_ds = train_ds.map(lambda a, b: mixup(a, b, args.batch_size), num_parallel_calls=AUTO)</code>. This augmentation takes a whole batch as input. To augment single images you'd move the augmentation before the <code>ds.batch</code> step.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 767756,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-03-10T04:07:09.283000",
      "content": "<p>Don't do inference with CPU for 3rd notebooks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 766691,
      "author_name": "Anindya Sanyal",
      "author_url": "",
      "post_date": "2020-03-08T15:14:37.803000",
      "content": "<p>Very Nice!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 766346,
      "author_name": "rick",
      "author_url": "",
      "post_date": "2020-03-08T03:22:49.003000",
      "content": "<p>Thank you for sharing! We can dramatically reduce training time. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 766076,
      "author_name": "Rohit Kumar",
      "author_url": "",
      "post_date": "2020-03-07T16:20:45.627000",
      "content": "<p>its nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 765734,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-03-07T02:16:59.537000",
      "content": "<p>I was thinking to try on TPU in this competition but couldn't get enough time. I have tried CapsNet though, but no luck.  Anyway, thanks for sharing, amazing work indeed. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 772154,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-15T05:20:12.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 770262,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-12T18:42:23.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 765471,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2020-03-06T16:55:52.523000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 765315,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2020-03-06T13:34:38.600000",
      "content": "<p>Many thanks! :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 765225,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2020-03-06T11:55:25.900000",
      "content": "<p>Great stuff as usual! Thank you!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 770077,
      "author_name": "hopeful1213",
      "author_url": "",
      "post_date": "2020-03-12T14:43:47.153000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 769979,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-12T13:05:32.353000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 768647,
      "author_name": "Rajesh Kumar",
      "author_url": "",
      "post_date": "2020-03-11T03:57:17.987000",
      "content": "<p>Thanks for sharing. Will try TPU.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 768051,
      "author_name": "Yuankun Liu",
      "author_url": "",
      "post_date": "2020-03-10T12:02:18.583000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767856,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-10T07:09:15.003000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "765194": "Hi everyone,\n\nThanks to Kaggle we can now use super fast TPUs for free! I wanted to apply them in this challenge so I created the following short series of Notebooks:\n\n- [1. Create TFRecords](https://www.kaggle.com/seesee/1-create-tfrecords)\n\nThis is a two-step process. The images are first converted to `.png` format and then converted to [`TFRecords`](https://www.tensorflow.org/tutorials/load_data/tfrecord)(thanks to @mgornergoogle and @rsmits for the code). Creating `TFRecords` is recommended for datasets that don't fit into memory. Note that you can skip this step as I already created the following dataset: https://www.kaggle.com/seesee/bengali-tfrecords-v010\n\n- [2. Train](https://www.kaggle.com/seesee/2-train)\n\nEfficientnet-B3 is trained for 25 epochs. You can easily switch the `backbone` thanks to the [efficientnet library](https://github.com/qubvel/efficientnet). There are 2 parts that are not \"vanilla\": I am using `mixup` augmentation and I train with `softmax` and then decode the predictions. There is a ton that can be improved. For example, you can try [these weights](https://www.kaggle.com/c/bengaliai-cv19/discussion/132894), a non-constant learning rate schedule, multiple folds, etc. The training takes less than 2 hours and I get 0.9708 on LB.\n\n- [3. Submit](https://www.kaggle.com/seesee/3-submit)\n\nTPU Notebooks are not allowed for submission so I created a new Notebook with GPU accelerator. The inference takes ~30 minutes.\n\nI hope this makes it easier to work with TPUs on Kaggle!",
    "770736": "Many thanks, I need a support on TPU for my research!",
    "766631": "Have anyone tried to train on colab TPU? I ran out my TPU quora this week, and then I started to work on reproduce this on google colab TPU, but keep failing....\nI still encounter this issue even though it is said fixed. \nhttps://github.com/googlecolab/colabtools/issues/808",
    "766054": "Hi! Would you explain what is \"tuple_map\" in submit, and how did you get it?\nThank you",
    "765771": "@seesee  Thanks for sharing this is really helpful!\n\nWhen I tried my idea based on your kernel, sometimes the kernel was suddenly killed and there was no error message.\nI ran the code with GPU and it ran successfully without any problem.\nI experienced almost same situation when I used pytorch xla.\n\nSo, my question is, is there any better way to debug code with TPU??\n\nThanks in advance!",
    "765706": "This is amazing! Thanks for the great work.",
    "765434": "@seesee, thanks so much for sharing. I hardly use my TPU quota and you just gave me a reason to experiment.",
    "765219": "Love you！brother goose🤓 ",
    "765215": "thx for the contribution. I wonder what's your motivation to use `one head + softmax` here. Any idea how it's going to perform on unseen combinations?",
    "765666": "just Wow many thanks @seesee , keep the great work",
    "770146": "Good day! When augmentation should be applied - to the original data, or somehow after they are TFR conversion?",
    "767756": "Don't do inference with CPU for 3rd notebooks",
    "766691": "Very Nice!",
    "766346": "Thank you for sharing! We can dramatically reduce training time. ",
    "766076": "its nice\n",
    "765734": "I was thinking to try on TPU in this competition but couldn't get enough time. I have tried CapsNet though, but no luck.  Anyway, thanks for sharing, amazing work indeed. ",
    "772154": "",
    "770262": "",
    "765471": "Thanks for sharing!",
    "765315": "Many thanks! :-)",
    "765225": "Great stuff as usual! Thank you!",
    "770077": "Thanks for sharing!",
    "769979": "Thanks for sharing",
    "768647": "Thanks for sharing. Will try TPU.",
    "768051": "Thanks for sharing!",
    "767856": "Thanks for sharing!"
  }
}