{
  "id": 143121,
  "title": "[QUESTION] is inference running on all of 8 cores with TPU?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143121",
  "author_name": "",
  "post_date": "2020-04-13T22:38:29.762051400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I feel inference speed on TPU is slower compared with the boost when training.\nMy setup is pretty similar with public kernel such as <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta/output\">this</a>, does anyone feel same thing? Or does anyone know the reason why?\nI'm not sure but I'm wondering whether inference is running on all of 8 cores or not.\nAny advice is helpful!\nThanks,</p>",
  "messages": [
    {
      "id": "806628",
      "postDate": "04/13/2020 22:38:29",
      "content": "<p>I feel inference speed on TPU is slower compared with the boost when training.\nMy setup is pretty similar with public kernel such as <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta/output\">this</a>, does anyone feel same thing? Or does anyone know the reason why?\nI'm not sure but I'm wondering whether inference is running on all of 8 cores or not.\nAny advice is helpful!\nThanks,</p>",
      "rawMarkdown": "I feel inference speed on TPU is slower compared with the boost when training.\nMy setup is pretty similar with public kernel such as [this](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta/output), does anyone feel same thing? Or does anyone know the reason why?\nI'm not sure but I'm wondering whether inference is running on all of 8 cores or not.\nAny advice is helpful!\nThanks,",
      "votes": null
    },
    {
      "id": "809450",
      "postDate": "04/16/2020 07:55:37",
      "content": "<p>Hi Camaro, my guess is that at the inference time, TPU workload is much lesser than training time, so the overhead comeback to our data pipeline.</p>\n\n<p>BTW, in the dashboard on the top-right of the notebook, we can see the TPU workload in %</p>",
      "rawMarkdown": "Hi Camaro, my guess is that at the inference time, TPU workload is much lesser than training time, so the overhead comeback to our data pipeline.\n\nBTW, in the dashboard on the top-right of the notebook, we can see the TPU workload in %",
      "votes": null
    },
    {
      "id": "809457",
      "postDate": "04/16/2020 07:57:41",
      "content": "<p>Also you do <code>.cpu()</code>; that's painfully slower;</p>",
      "rawMarkdown": "Also you do `.cpu()`; that's painfully slower;",
      "votes": null
    },
    {
      "id": "809742",
      "postDate": "04/16/2020 12:49:48",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> Thanks, possibly. I'll check it:)</p>",
      "rawMarkdown": "ratthachat Thanks, possibly. I'll check it:)",
      "votes": null
    },
    {
      "id": "809745",
      "postDate": "04/16/2020 12:52:36",
      "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> I believe I didn't, if so it takes literally forever😸 </p>",
      "rawMarkdown": "adityaecdrid I believe I didn't, if so it takes literally forever😸",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 809450,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "04/16/2020 07:55:37",
      "content": "<p>Hi Camaro, my guess is that at the inference time, TPU workload is much lesser than training time, so the overhead comeback to our data pipeline.</p>\n\n<p>BTW, in the dashboard on the top-right of the notebook, we can see the TPU workload in %</p>",
      "votes": null,
      "replies": [
        {
          "id": 809457,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "04/16/2020 07:57:41",
          "content": "<p>Also you do <code>.cpu()</code>; that's painfully slower;</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809742,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/16/2020 12:49:48",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> Thanks, possibly. I'll check it:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809745,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/16/2020 12:52:36",
          "content": "<p><a href=\"/adityaecdrid\">@adityaecdrid</a> I believe I didn't, if so it takes literally forever😸 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "806628": "I feel inference speed on TPU is slower compared with the boost when training.\nMy setup is pretty similar with public kernel such as [this](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta/output), does anyone feel same thing? Or does anyone know the reason why?\nI'm not sure but I'm wondering whether inference is running on all of 8 cores or not.\nAny advice is helpful!\nThanks,",
    "809450": "Hi Camaro, my guess is that at the inference time, TPU workload is much lesser than training time, so the overhead comeback to our data pipeline.\n\nBTW, in the dashboard on the top-right of the notebook, we can see the TPU workload in %",
    "809457": "Also you do `.cpu()`; that's painfully slower;",
    "809742": "ratthachat Thanks, possibly. I'll check it:)",
    "809745": "adityaecdrid I believe I didn't, if so it takes literally forever😸"
  },
  "source": "meta"
}