{
  "id": 511785,
  "title": "RuntimeError: Mixing different tf.distribute.Strategy objects",
  "url": "/competitions/tpu-getting-started/discussion/511785",
  "author_name": "Kawchar Husain",
  "post_date": "2024-06-12T05:49:56.060000",
  "votes": 4,
  "comment_count": 3,
  "views": null,
  "content": "<p>what is the problem in model training?</p>\n<p>I'm following this <a href=\"https://www.kaggle.com/code/ryanholbrook/create-your-first-submission\" target=\"_blank\">tutorial</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12376195%2F4197cb74d802440a426325edeb7397fa%2FScreenshot%202024-06-12%20at%2011.45.11AM.png?generation=1718171172835181&amp;alt=media\"></p>",
  "messages": [
    {
      "id": 2867852,
      "postDate": "2024-06-12T05:49:56.060Z",
      "content": "<p>what is the problem in model training?</p>\n<p>I'm following this <a href=\"https://www.kaggle.com/code/ryanholbrook/create-your-first-submission\" target=\"_blank\">tutorial</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12376195%2F4197cb74d802440a426325edeb7397fa%2FScreenshot%202024-06-12%20at%2011.45.11AM.png?generation=1718171172835181&amp;alt=media\"></p>",
      "rawMarkdown": "what is the problem in model training?\n\nI'm following this [tutorial](https://www.kaggle.com/code/ryanholbrook/create-your-first-submission)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12376195%2F4197cb74d802440a426325edeb7397fa%2FScreenshot%202024-06-12%20at%2011.45.11AM.png?generation=1718171172835181&alt=media)",
      "votes": 4
    },
    {
      "id": 2997640,
      "postDate": "2024-09-24T18:24:01.673Z",
      "content": "<p>Put model compiling and fitting in the strategy.scope() (in the cell where \"history.fit=model.fit\" is defined) and don't change other cells:</p>\n<pre><code> strategy.scope():    \n    pretrained_model = tf.keras.applications.xception.Xception(\n        weights=, \n        include_top=,\n        input_shape=[*IMAGE_SIZE, ]\n    )\n    pretrained_model.trainable =  \n\n    model = tf.keras.Sequential([\n        pretrained_model,\n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dense(, activation=)\n    ])\n\n    model.(\n        optimizer=,\n        loss = ,\n        metrics=[]\n    )\n\n    history= model.fit(training_dataset, \n              steps_per_epoch=STEPS_PER_EPOCH, \n              epochs=EPOCHS, \n              validation_data=validation_dataset)\n</code></pre>\n<p>It should help, but you certainly will face with a mistake in the \"confusion_matrix\" block. I haven't found any way to fix it and just used another notebook.</p>",
      "rawMarkdown": "Put model compiling and fitting in the strategy.scope() (in the cell where \"history.fit=model.fit\" is defined) and don't change other cells:\n\n```python\nwith strategy.scope():    \n    pretrained_model = tf.keras.applications.xception.Xception(\n        weights='imagenet', \n        include_top=False,\n        input_shape=[*IMAGE_SIZE, 3]\n    )\n    pretrained_model.trainable = False # tramsfer learning\n\n    model = tf.keras.Sequential([\n        pretrained_model,\n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dense(104, activation='softmax')\n    ])\n        \n    model.compile(\n        optimizer='adam',\n        loss = 'sparse_categorical_crossentropy',\n        metrics=['sparse_categorical_accuracy']\n    )\n\n    history= model.fit(training_dataset, \n              steps_per_epoch=STEPS_PER_EPOCH, \n              epochs=EPOCHS, \n              validation_data=validation_dataset)\n```\n\nIt should help, but you certainly will face with a mistake in the \"confusion_matrix\" block. I haven't found any way to fix it and just used another notebook.",
      "votes": 2
    },
    {
      "id": 3157724,
      "postDate": "2025-03-23T18:38:33.343Z",
      "content": "<p>For the recent TF changes 2.18.0 I have created a fix of \"a simple Petals 2.2 TF Notebook\" here <a href=\"https://www.kaggle.com/code/sergeivasilenko/a-simple-petals-tf-2-18-notebook-fixed\" target=\"_blank\">https://www.kaggle.com/code/sergeivasilenko/a-simple-petals-tf-2-18-notebook-fixed</a>. There are two major fixes: </p>\n<p>The first the initialisation of TPUStrategy is not experimental anymore and it is required the following code: </p>\n<pre><code>\n:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=) \n    (, tpu.master())\n ValueError:\n    ()\n    tpu = \n\n tpu:\n    strategy = tf.distribute.TPUStrategy(tpu)\n:\n    strategy = tf.distribute.get_strategy() \n\n(, strategy.num_replicas_in_sync)\n</code></pre>\n<p>The second fix is covered by YaEmil in his answer <a href=\"https://www.kaggle.com/competitions/tpu-getting-started/discussion/511785#2997640\" target=\"_blank\">https://www.kaggle.com/competitions/tpu-getting-started/discussion/511785#2997640</a> that compiling and fitting should be inside of <code>strategy.scope()</code></p>",
      "rawMarkdown": "For the recent TF changes 2.18.0 I have created a fix of \"a simple Petals 2.2 TF Notebook\" here https://www.kaggle.com/code/sergeivasilenko/a-simple-petals-tf-2-18-notebook-fixed. There are two major fixes: \n\nThe first the initialisation of TPUStrategy is not experimental anymore and it is required the following code: \n\n```\n# Detect hardware, return appropriate distribution strategy\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu='local') # set tpu is local as it should be available in the VM\n    print('✅ Running on TPU ', tpu.master())\nexcept ValueError:\n    print('❌ Using CPU/GPU')\n    tpu = None\n\nif tpu:\n    strategy = tf.distribute.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() # default distribution strategy in Tensorflow. Works on CPU and single GPU.\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```\n\nThe second fix is covered by YaEmil in his answer https://www.kaggle.com/competitions/tpu-getting-started/discussion/511785#2997640 that compiling and fitting should be inside of `strategy.scope()`"
    },
    {
      "id": 2986357,
      "postDate": "2024-09-11T15:25:35.333Z",
      "content": "<p>Hi, I'm having the same problem. Did you ever find the solution?</p>",
      "rawMarkdown": "Hi, I'm having the same problem. Did you ever find the solution?"
    }
  ],
  "comments": [
    {
      "id": 2997640,
      "author_name": "YaEmil",
      "author_url": "",
      "post_date": "2024-09-24T18:24:01.673000",
      "content": "<p>Put model compiling and fitting in the strategy.scope() (in the cell where \"history.fit=model.fit\" is defined) and don't change other cells:</p>\n<pre><code> strategy.scope():    \n    pretrained_model = tf.keras.applications.xception.Xception(\n        weights=, \n        include_top=,\n        input_shape=[*IMAGE_SIZE, ]\n    )\n    pretrained_model.trainable =  \n\n    model = tf.keras.Sequential([\n        pretrained_model,\n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dense(, activation=)\n    ])\n\n    model.(\n        optimizer=,\n        loss = ,\n        metrics=[]\n    )\n\n    history= model.fit(training_dataset, \n              steps_per_epoch=STEPS_PER_EPOCH, \n              epochs=EPOCHS, \n              validation_data=validation_dataset)\n</code></pre>\n<p>It should help, but you certainly will face with a mistake in the \"confusion_matrix\" block. I haven't found any way to fix it and just used another notebook.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3157724,
      "author_name": "Sergei Vasilenko",
      "author_url": "",
      "post_date": "2025-03-23T18:38:33.343000",
      "content": "<p>For the recent TF changes 2.18.0 I have created a fix of \"a simple Petals 2.2 TF Notebook\" here <a href=\"https://www.kaggle.com/code/sergeivasilenko/a-simple-petals-tf-2-18-notebook-fixed\" target=\"_blank\">https://www.kaggle.com/code/sergeivasilenko/a-simple-petals-tf-2-18-notebook-fixed</a>. There are two major fixes: </p>\n<p>The first the initialisation of TPUStrategy is not experimental anymore and it is required the following code: </p>\n<pre><code>\n:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu=) \n    (, tpu.master())\n ValueError:\n    ()\n    tpu = \n\n tpu:\n    strategy = tf.distribute.TPUStrategy(tpu)\n:\n    strategy = tf.distribute.get_strategy() \n\n(, strategy.num_replicas_in_sync)\n</code></pre>\n<p>The second fix is covered by YaEmil in his answer <a href=\"https://www.kaggle.com/competitions/tpu-getting-started/discussion/511785#2997640\" target=\"_blank\">https://www.kaggle.com/competitions/tpu-getting-started/discussion/511785#2997640</a> that compiling and fitting should be inside of <code>strategy.scope()</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2986357,
      "author_name": "Romeo Angasa",
      "author_url": "",
      "post_date": "2024-09-11T15:25:35.333000",
      "content": "<p>Hi, I'm having the same problem. Did you ever find the solution?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2867852": "what is the problem in model training?\n\nI'm following this [tutorial](https://www.kaggle.com/code/ryanholbrook/create-your-first-submission)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12376195%2F4197cb74d802440a426325edeb7397fa%2FScreenshot%202024-06-12%20at%2011.45.11AM.png?generation=1718171172835181&alt=media)",
    "2997640": "Put model compiling and fitting in the strategy.scope() (in the cell where \"history.fit=model.fit\" is defined) and don't change other cells:\n\n```python\nwith strategy.scope():    \n    pretrained_model = tf.keras.applications.xception.Xception(\n        weights='imagenet', \n        include_top=False,\n        input_shape=[*IMAGE_SIZE, 3]\n    )\n    pretrained_model.trainable = False # tramsfer learning\n\n    model = tf.keras.Sequential([\n        pretrained_model,\n        tf.keras.layers.GlobalAveragePooling2D(),\n        tf.keras.layers.Dense(104, activation='softmax')\n    ])\n        \n    model.compile(\n        optimizer='adam',\n        loss = 'sparse_categorical_crossentropy',\n        metrics=['sparse_categorical_accuracy']\n    )\n\n    history= model.fit(training_dataset, \n              steps_per_epoch=STEPS_PER_EPOCH, \n              epochs=EPOCHS, \n              validation_data=validation_dataset)\n```\n\nIt should help, but you certainly will face with a mistake in the \"confusion_matrix\" block. I haven't found any way to fix it and just used another notebook.",
    "3157724": "For the recent TF changes 2.18.0 I have created a fix of \"a simple Petals 2.2 TF Notebook\" here https://www.kaggle.com/code/sergeivasilenko/a-simple-petals-tf-2-18-notebook-fixed. There are two major fixes: \n\nThe first the initialisation of TPUStrategy is not experimental anymore and it is required the following code: \n\n```\n# Detect hardware, return appropriate distribution strategy\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect(tpu='local') # set tpu is local as it should be available in the VM\n    print('✅ Running on TPU ', tpu.master())\nexcept ValueError:\n    print('❌ Using CPU/GPU')\n    tpu = None\n\nif tpu:\n    strategy = tf.distribute.TPUStrategy(tpu)\nelse:\n    strategy = tf.distribute.get_strategy() # default distribution strategy in Tensorflow. Works on CPU and single GPU.\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n```\n\nThe second fix is covered by YaEmil in his answer https://www.kaggle.com/competitions/tpu-getting-started/discussion/511785#2997640 that compiling and fitting should be inside of `strategy.scope()`",
    "2986357": "Hi, I'm having the same problem. Did you ever find the solution?"
  }
}