{
  "id": 271219,
  "title": "TPU convergence problems",
  "url": "/competitions/landmark-recognition-2021/discussion/271219",
  "author_name": "",
  "post_date": "2021-09-09T07:35:56.351038700Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone!</p>\n<p>Have someone faced problems with TPU convergence in this competition?</p>\n<p>I have tried to train model with this configuration</p>\n<p><code>x = EfficientNetB3(weights='imagenet', include_top=False)(x)\n  x = tf.keras.layers.GlobalAveragePooling2D()(x)\n  x = tf.keras.layers.Dense(512)(x)\n  x = tf.keras.layers.BatchNormalization()(x)</code></p>\n<ul>\n<li>Optimizer: Adam</li>\n<li>Loss: CrossEntropyLoss</li>\n</ul>\n<p>using all clean GLRDv2 and also it's subsets, but almost every time the model trained on TPU converges to predicting only the most frequent class during the first epoch.<br>\nOn the other side, when training on GPU everything is okay and it doesn't go in such a pitfall, so I assume there is some problem in my TPU initialisation or something.</p>\n<p>I will be grateful if someone can explain where is the problem.</p>",
  "messages": [
    {
      "id": "1507427",
      "postDate": "09/09/2021 07:35:56",
      "content": "<p>Hello everyone!</p>\n<p>Have someone faced problems with TPU convergence in this competition?</p>\n<p>I have tried to train model with this configuration</p>\n<p><code>x = EfficientNetB3(weights='imagenet', include_top=False)(x)\n  x = tf.keras.layers.GlobalAveragePooling2D()(x)\n  x = tf.keras.layers.Dense(512)(x)\n  x = tf.keras.layers.BatchNormalization()(x)</code></p>\n<ul>\n<li>Optimizer: Adam</li>\n<li>Loss: CrossEntropyLoss</li>\n</ul>\n<p>using all clean GLRDv2 and also it's subsets, but almost every time the model trained on TPU converges to predicting only the most frequent class during the first epoch.<br>\nOn the other side, when training on GPU everything is okay and it doesn't go in such a pitfall, so I assume there is some problem in my TPU initialisation or something.</p>\n<p>I will be grateful if someone can explain where is the problem.</p>",
      "rawMarkdown": "Hello everyone!\n\nHave someone faced problems with TPU convergence in this competition?\n\nI have tried to train model with this configuration\n\n`x = EfficientNetB3(weights='imagenet', include_top=False)(x)\n  x = tf.keras.layers.GlobalAveragePooling2D()(x)\n  x = tf.keras.layers.Dense(512)(x)\n  x = tf.keras.layers.BatchNormalization()(x)`\n\n+ Optimizer: Adam\n+ Loss: CrossEntropyLoss\n\nusing all clean GLRDv2 and also it's subsets, but almost every time the model trained on TPU converges to predicting only the most frequent class during the first epoch.\nOn the other side, when training on GPU everything is okay and it doesn't go in such a pitfall, so I assume there is some problem in my TPU initialisation or something.\n \nI will be grateful if someone can explain where is the problem.",
      "votes": null
    },
    {
      "id": "1507449",
      "postDate": "09/09/2021 08:03:06",
      "content": "<p>What batch size and learning rate do you use? In general, a TPU requires an 8 times higher batch size and learning rate due to the 8 computing units.</p>",
      "rawMarkdown": "What batch size and learning rate do you use? In general, a TPU requires an 8 times higher batch size and learning rate due to the 8 computing units.",
      "votes": null
    },
    {
      "id": "1507568",
      "postDate": "09/09/2021 10:27:45",
      "content": "<p>Hello!</p>\n<p>Yes, I take this into account and multiply my batch size by 8 compared to GPU batch size.<br>\nIn fact, for 256x256 resolution images I use 64 * 8 batch size and learning rate 1e-4.<br>\nI also tried to increase my lr by a factor of 8, but that didn't help either</p>",
      "rawMarkdown": "Hello!\n\nYes, I take this into account and multiply my batch size by 8 compared to GPU batch size.\nIn fact, for 256x256 resolution images I use 64 * 8 batch size and learning rate 1e-4.\nI also tried to increase my lr by a factor of 8, but that didn't help either",
      "votes": null
    },
    {
      "id": "1512818",
      "postDate": "09/14/2021 15:56:44",
      "content": "<p>I finally solved the problem by setting <strong>determenistic = True</strong> for tensorflow dataset options.</p>\n<p>Precisely, I have changed this block:</p>\n<p><code>ignore_order = tf.data.Options()\n  ignore_order.experimental_deterministic = False\n  dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.with_options(ignore_order)</code></p>\n<p>to:</p>\n<p><code>dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.map(parse_fn, num_parallel_calls=AUTOTUNE)</code></p>\n<p>After that the model begin to learn properly on TPU.</p>",
      "rawMarkdown": "I finally solved the problem by setting **determenistic = True** for tensorflow dataset options.\n\nPrecisely, I have changed this block:\n\n`ignore_order = tf.data.Options()\n  ignore_order.experimental_deterministic = False\n  dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.with_options(ignore_order)`\n\nto:\n\n`dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.map(parse_fn, num_parallel_calls=AUTOTUNE)`\n\nAfter that the model begin to learn properly on TPU.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1507449,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "09/09/2021 08:03:06",
      "content": "<p>What batch size and learning rate do you use? In general, a TPU requires an 8 times higher batch size and learning rate due to the 8 computing units.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1507568,
          "author_name": "andreygalichin",
          "author_url": "",
          "post_date": "09/09/2021 10:27:45",
          "content": "<p>Hello!</p>\n<p>Yes, I take this into account and multiply my batch size by 8 compared to GPU batch size.<br>\nIn fact, for 256x256 resolution images I use 64 * 8 batch size and learning rate 1e-4.<br>\nI also tried to increase my lr by a factor of 8, but that didn't help either</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1512818,
      "author_name": "andreygalichin",
      "author_url": "",
      "post_date": "09/14/2021 15:56:44",
      "content": "<p>I finally solved the problem by setting <strong>determenistic = True</strong> for tensorflow dataset options.</p>\n<p>Precisely, I have changed this block:</p>\n<p><code>ignore_order = tf.data.Options()\n  ignore_order.experimental_deterministic = False\n  dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.with_options(ignore_order)</code></p>\n<p>to:</p>\n<p><code>dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.map(parse_fn, num_parallel_calls=AUTOTUNE)</code></p>\n<p>After that the model begin to learn properly on TPU.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1507427": "Hello everyone!\n\nHave someone faced problems with TPU convergence in this competition?\n\nI have tried to train model with this configuration\n\n`x = EfficientNetB3(weights='imagenet', include_top=False)(x)\n  x = tf.keras.layers.GlobalAveragePooling2D()(x)\n  x = tf.keras.layers.Dense(512)(x)\n  x = tf.keras.layers.BatchNormalization()(x)`\n\n+ Optimizer: Adam\n+ Loss: CrossEntropyLoss\n\nusing all clean GLRDv2 and also it's subsets, but almost every time the model trained on TPU converges to predicting only the most frequent class during the first epoch.\nOn the other side, when training on GPU everything is okay and it doesn't go in such a pitfall, so I assume there is some problem in my TPU initialisation or something.\n \nI will be grateful if someone can explain where is the problem.",
    "1507449": "What batch size and learning rate do you use? In general, a TPU requires an 8 times higher batch size and learning rate due to the 8 computing units.",
    "1507568": "Hello!\n\nYes, I take this into account and multiply my batch size by 8 compared to GPU batch size.\nIn fact, for 256x256 resolution images I use 64 * 8 batch size and learning rate 1e-4.\nI also tried to increase my lr by a factor of 8, but that didn't help either",
    "1512818": "I finally solved the problem by setting **determenistic = True** for tensorflow dataset options.\n\nPrecisely, I have changed this block:\n\n`ignore_order = tf.data.Options()\n  ignore_order.experimental_deterministic = False\n  dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.with_options(ignore_order)`\n\nto:\n\n`dataset = tf.data.TFRecordDataset(paths, num_parallel_reads=AUTOTUNE)\n  dataset = dataset.map(parse_fn, num_parallel_calls=AUTOTUNE)`\n\nAfter that the model begin to learn properly on TPU."
  },
  "source": "meta"
}