{
  "id": 141413,
  "title": "number of data samples is not divisible by batch size when I run on TPU",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141413",
  "author_name": "",
  "post_date": "2020-04-06T00:02:01.397520800Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I keep getting this error : ValueError: The number of samples 8305 is not divisible by batch size 16.</p>\n\n<p>I tried only a small sample of the data but I keep on getting this error, anyone encountered this problem?</p>",
  "messages": [
    {
      "id": "798869",
      "postDate": "04/06/2020 00:02:01",
      "content": "<p>I keep getting this error : ValueError: The number of samples 8305 is not divisible by batch size 16.</p>\n\n<p>I tried only a small sample of the data but I keep on getting this error, anyone encountered this problem?</p>",
      "rawMarkdown": "I keep getting this error : ValueError: The number of samples 8305 is not divisible by batch size 16.\n\nI tried only a small sample of the data but I keep on getting this error, anyone encountered this problem?",
      "votes": null
    },
    {
      "id": "798872",
      "postDate": "04/06/2020 00:10:53",
      "content": "<p>Make it divisible then, as simple as that;</p>",
      "rawMarkdown": "Make it divisible then, as simple as that;",
      "votes": null
    },
    {
      "id": "798890",
      "postDate": "04/06/2020 00:44:20",
      "content": "<p>Would something like this work to limit the number of calls?\n```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\ntf.config.experimental_connect_to_cluster(tpu)\ntf.tpu.experimental.initialize_tpu_system(tpu)\nstrategy = tf.distribute.experimental.TPUStrategy(tpu)</p>\n\n<p>BATCH_SIZE = 16 * strategy.num_replicas_in_sync\nTRAIN_DATA_LENGTH = 223549  # rows\nSTEPS_PER_EPOCH = TRAIN_DATA_LENGTH // BATCH_SIZE</p>\n\n<p>history = multilingual_bert.fit(\n    english_train_dataset,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    epochs=3,\n    verbose=1,\n    validation_data=nonenglish_val_datasets[\"Combined\"],\n)\n<code>``\nSpecifically, I'd expect the</code>steps_per_epoch<code>argument will limit the number of steps.  You might need to call</code>dataset.repeat()` on the dataset in case there's an overrun, but the integer divide in STEPS_PER_EPOCH may make repeat() unnecessary.</p>\n\n<p>I'm a bit of a noob, so maybe I'm wrong...</p>",
      "rawMarkdown": "Would something like this work to limit the number of calls?\n```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\ntf.config.experimental_connect_to_cluster(tpu)\ntf.tpu.experimental.initialize_tpu_system(tpu)\nstrategy = tf.distribute.experimental.TPUStrategy(tpu)\n\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\nTRAIN_DATA_LENGTH = 223549  # rows\nSTEPS_PER_EPOCH = TRAIN_DATA_LENGTH // BATCH_SIZE\n\nhistory = multilingual_bert.fit(\n    english_train_dataset,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    epochs=3,\n    verbose=1,\n    validation_data=nonenglish_val_datasets[\"Combined\"],\n)\n```\nSpecifically, I'd expect the `steps_per_epoch` argument will limit the number of steps.  You might need to call `dataset.repeat()` on the dataset in case there's an overrun, but the integer divide in STEPS_PER_EPOCH may make repeat() unnecessary.\n\nI'm a bit of a noob, so maybe I'm wrong...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 798872,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "04/06/2020 00:10:53",
      "content": "<p>Make it divisible then, as simple as that;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 798890,
      "author_name": "matthalenka",
      "author_url": "",
      "post_date": "04/06/2020 00:44:20",
      "content": "<p>Would something like this work to limit the number of calls?\n```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\ntf.config.experimental_connect_to_cluster(tpu)\ntf.tpu.experimental.initialize_tpu_system(tpu)\nstrategy = tf.distribute.experimental.TPUStrategy(tpu)</p>\n\n<p>BATCH_SIZE = 16 * strategy.num_replicas_in_sync\nTRAIN_DATA_LENGTH = 223549  # rows\nSTEPS_PER_EPOCH = TRAIN_DATA_LENGTH // BATCH_SIZE</p>\n\n<p>history = multilingual_bert.fit(\n    english_train_dataset,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    epochs=3,\n    verbose=1,\n    validation_data=nonenglish_val_datasets[\"Combined\"],\n)\n<code>``\nSpecifically, I'd expect the</code>steps_per_epoch<code>argument will limit the number of steps.  You might need to call</code>dataset.repeat()` on the dataset in case there's an overrun, but the integer divide in STEPS_PER_EPOCH may make repeat() unnecessary.</p>\n\n<p>I'm a bit of a noob, so maybe I'm wrong...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "798869": "I keep getting this error : ValueError: The number of samples 8305 is not divisible by batch size 16.\n\nI tried only a small sample of the data but I keep on getting this error, anyone encountered this problem?",
    "798872": "Make it divisible then, as simple as that;",
    "798890": "Would something like this work to limit the number of calls?\n```\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver()  # TPU detection\ntf.config.experimental_connect_to_cluster(tpu)\ntf.tpu.experimental.initialize_tpu_system(tpu)\nstrategy = tf.distribute.experimental.TPUStrategy(tpu)\n\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\nTRAIN_DATA_LENGTH = 223549  # rows\nSTEPS_PER_EPOCH = TRAIN_DATA_LENGTH // BATCH_SIZE\n\nhistory = multilingual_bert.fit(\n    english_train_dataset,\n    steps_per_epoch=STEPS_PER_EPOCH,\n    epochs=3,\n    verbose=1,\n    validation_data=nonenglish_val_datasets[\"Combined\"],\n)\n```\nSpecifically, I'd expect the `steps_per_epoch` argument will limit the number of steps.  You might need to call `dataset.repeat()` on the dataset in case there's an overrun, but the integer divide in STEPS_PER_EPOCH may make repeat() unnecessary.\n\nI'm a bit of a noob, so maybe I'm wrong..."
  },
  "source": "meta"
}