{
  "id": 233389,
  "title": "Distributed Training and Inference on TPU",
  "url": "/competitions/bms-molecular-translation/discussion/233389",
  "author_name": "Jude TCHAYE",
  "post_date": "2021-04-19T01:10:49.897000",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello!<br>\nI just released some resources on how to conduct a distributed training on TPU. I started by converting the entire dataset into blocks of tfrecord files(around 100MB each block). The conversion notebook is at <a href=\"https://www.kaggle.com/tchaye59/mt-tfrecord-custom-vocab\" target=\"_blank\">https://www.kaggle.com/tchaye59/mt-tfrecord-custom-vocab</a> and the resulting dataset at <a href=\"https://www.kaggle.com/tchaye59/mtcustomvocabimg\" target=\"_blank\">https://www.kaggle.com/tchaye59/mtcustomvocabimg</a>. The Encoder, Decoder with Attention Mechanism was used for the training: <a href=\"https://www.kaggle.com/tchaye59/mt-fast-distributed-training-tpu\" target=\"_blank\">https://www.kaggle.com/tchaye59/mt-fast-distributed-training-tpu</a>. The training pipeline is only tested on TPU, and the global training bach size is 512*8=4096. (It cloud run on CPU or GPU with reduced batch size)</p>",
  "messages": [
    {
      "id": 1277594,
      "postDate": "2021-04-19T01:10:49.897Z",
      "content": "<p>Hello!<br>\nI just released some resources on how to conduct a distributed training on TPU. I started by converting the entire dataset into blocks of tfrecord files(around 100MB each block). The conversion notebook is at <a href=\"https://www.kaggle.com/tchaye59/mt-tfrecord-custom-vocab\" target=\"_blank\">https://www.kaggle.com/tchaye59/mt-tfrecord-custom-vocab</a> and the resulting dataset at <a href=\"https://www.kaggle.com/tchaye59/mtcustomvocabimg\" target=\"_blank\">https://www.kaggle.com/tchaye59/mtcustomvocabimg</a>. The Encoder, Decoder with Attention Mechanism was used for the training: <a href=\"https://www.kaggle.com/tchaye59/mt-fast-distributed-training-tpu\" target=\"_blank\">https://www.kaggle.com/tchaye59/mt-fast-distributed-training-tpu</a>. The training pipeline is only tested on TPU, and the global training bach size is 512*8=4096. (It cloud run on CPU or GPU with reduced batch size)</p>",
      "rawMarkdown": "Hello!\nI just released some resources on how to conduct a distributed training on TPU. I started by converting the entire dataset into blocks of tfrecord files(around 100MB each block). The conversion notebook is at https://www.kaggle.com/tchaye59/mt-tfrecord-custom-vocab and the resulting dataset at https://www.kaggle.com/tchaye59/mtcustomvocabimg. The Encoder, Decoder with Attention Mechanism was used for the training: https://www.kaggle.com/tchaye59/mt-fast-distributed-training-tpu. The training pipeline is only tested on TPU, and the global training bach size is 512*8=4096. (It cloud run on CPU or GPU with reduced batch size)",
      "votes": 11
    },
    {
      "id": 1281432,
      "postDate": "2021-04-22T23:51:36.840Z",
      "content": "<p>How long did it take to convert everything (train and test) to tfrecord?</p>",
      "rawMarkdown": "How long did it take to convert everything (train and test) to tfrecord?",
      "votes": 1,
      "replies": [
        {
          "id": 1281477,
          "postDate": "2021-04-23T01:57:54.663Z",
          "content": "<p>It took longtime around 7 hours for train and 5 hours for test. I first converted the train set. After I run the notebook a second time to only convert the test set and copy back the train dataset</p>",
          "rawMarkdown": "It took longtime around 7 hours for train and 5 hours for test. I first converted the train set. After I run the notebook a second time to only convert the test set and copy back the train dataset",
          "votes": 1
        },
        {
          "id": 1289421,
          "postDate": "2021-05-01T00:41:36.303Z",
          "content": "<p>Did you have any trouble loading 9GB back to Kaggle??  </p>",
          "rawMarkdown": "Did you have any trouble loading 9GB back to Kaggle??  ",
          "votes": 1
        },
        {
          "id": 1289445,
          "postDate": "2021-05-01T02:08:05.443Z",
          "content": "<p>No check training notebook</p>",
          "rawMarkdown": "No check training notebook"
        }
      ]
    },
    {
      "id": 1277604,
      "postDate": "2021-04-19T01:28:09.463Z",
      "content": "<p>good work!</p>\n<p>i think one day, there will be a transformer encoder+decoder TPU public notebook …. that will be killer<br>\nin the end, the minimum score for bronze could be below 1.5 to 1.7</p>",
      "rawMarkdown": "good work!\n\ni think one day, there will be a transformer encoder+decoder TPU public notebook .... that will be killer\nin the end, the minimum score for bronze could be below 1.5 to 1.7",
      "votes": 1,
      "replies": [
        {
          "id": 1281855,
          "postDate": "2021-04-23T11:16:48.087Z",
          "content": "<p>I'm working on transformer / TPU, but there are still some problems. :(</p>",
          "rawMarkdown": "I'm working on transformer / TPU, but there are still some problems. :("
        },
        {
          "id": 1281883,
          "postDate": "2021-04-23T11:37:58.677Z",
          "content": "<p>We can work together if you want 🧐</p>",
          "rawMarkdown": "We can work together if you want 🧐",
          "votes": 1
        },
        {
          "id": 1283663,
          "postDate": "2021-04-25T06:38:13.863Z",
          "content": "<p>i open-source the input=patch+coord version in <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>\n<p>i think this can run very fast in TPU</p>\n<p>basically, there is no image. It is pure seq-to-seq.</p>\n<p>the input token is image patch of dim=16x16 pixel</p>\n<p>if it can be run very fast, maybe one can train from scratch</p>",
          "rawMarkdown": "i open-source the input=patch+coord version in https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n\ni think this can run very fast in TPU\n\nbasically, there is no image. It is pure seq-to-seq.\n\nthe input token is image patch of dim=16x16 pixel\n\nif it can be run very fast, maybe one can train from scratch",
          "votes": 1
        }
      ]
    },
    {
      "id": 1283628,
      "postDate": "2021-04-25T05:38:47.420Z",
      "content": "<p>Looks very neat. I would like to use them as my study references, thank you <a href=\"https://www.kaggle.com/tchaye59\" target=\"_blank\">@tchaye59</a>.</p>",
      "rawMarkdown": "Looks very neat. I would like to use them as my study references, thank you @tchaye59.",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1281432,
      "author_name": "Matt Yates",
      "author_url": "",
      "post_date": "2021-04-22T23:51:36.840000",
      "content": "<p>How long did it take to convert everything (train and test) to tfrecord?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1281477,
          "author_name": "Jude TCHAYE",
          "author_url": "",
          "post_date": "2021-04-23T01:57:54.663000",
          "content": "<p>It took longtime around 7 hours for train and 5 hours for test. I first converted the train set. After I run the notebook a second time to only convert the test set and copy back the train dataset</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1289421,
          "author_name": "Matt Yates",
          "author_url": "",
          "post_date": "2021-05-01T00:41:36.303000",
          "content": "<p>Did you have any trouble loading 9GB back to Kaggle??  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1289445,
          "author_name": "Jude TCHAYE",
          "author_url": "",
          "post_date": "2021-05-01T02:08:05.443000",
          "content": "<p>No check training notebook</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1277604,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-19T01:28:09.463000",
      "content": "<p>good work!</p>\n<p>i think one day, there will be a transformer encoder+decoder TPU public notebook …. that will be killer<br>\nin the end, the minimum score for bronze could be below 1.5 to 1.7</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1281855,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-04-23T11:16:48.087000",
          "content": "<p>I'm working on transformer / TPU, but there are still some problems. :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1281883,
          "author_name": "Jude TCHAYE",
          "author_url": "",
          "post_date": "2021-04-23T11:37:58.677000",
          "content": "<p>We can work together if you want 🧐</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1283663,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-04-25T06:38:13.863000",
          "content": "<p>i open-source the input=patch+coord version in <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>\n<p>i think this can run very fast in TPU</p>\n<p>basically, there is no image. It is pure seq-to-seq.</p>\n<p>the input token is image patch of dim=16x16 pixel</p>\n<p>if it can be run very fast, maybe one can train from scratch</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1283628,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-25T05:38:47.420000",
      "content": "<p>Looks very neat. I would like to use them as my study references, thank you <a href=\"https://www.kaggle.com/tchaye59\" target=\"_blank\">@tchaye59</a>.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1277594": "Hello!\nI just released some resources on how to conduct a distributed training on TPU. I started by converting the entire dataset into blocks of tfrecord files(around 100MB each block). The conversion notebook is at https://www.kaggle.com/tchaye59/mt-tfrecord-custom-vocab and the resulting dataset at https://www.kaggle.com/tchaye59/mtcustomvocabimg. The Encoder, Decoder with Attention Mechanism was used for the training: https://www.kaggle.com/tchaye59/mt-fast-distributed-training-tpu. The training pipeline is only tested on TPU, and the global training bach size is 512*8=4096. (It cloud run on CPU or GPU with reduced batch size)",
    "1281432": "How long did it take to convert everything (train and test) to tfrecord?",
    "1277604": "good work!\n\ni think one day, there will be a transformer encoder+decoder TPU public notebook .... that will be killer\nin the end, the minimum score for bronze could be below 1.5 to 1.7",
    "1283628": "Looks very neat. I would like to use them as my study references, thank you @tchaye59."
  }
}