{
  "id": 159579,
  "title": "MXU usage / how to get the most out of TPUs?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/159579",
  "author_name": "AgentAuers",
  "post_date": "2020-06-17T21:57:05.748000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The MXU is a specialized processing unit of the TPUs. They are used to perform very fast matrix operations like multiplications, which is the main computational work when training neural nets. The TPUs also have a more generalized CPU like processing unit, used to fetch and prepare data, communicate with the host, and coordinating the processing \"workflow\".</p>\n\n<p>The documentation says you should use \n- big batch sizes\n- multiple tfrec files\n- turn of the deterministic order when reading the files. (experimental_deterministic = False)\n- tf.dataset.cache()</p>\n\n<p>Most notebooks use above points, but I have not found one, which gets the MXU over 70-80% usage. </p>\n\n<p>What must be done to get the most out of the mxu or whole TPU board? Huge/deep nets? very large input/image sizes? I mean, it seams pretty impossible to saturate the MXUs of a single TPU board. What do the big companies do when learing with complete TPU-Pods?</p>\n\n<p>While digging around this, I created following notebook (<a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">Link</a>), which combines seven EfficientNets (B0-B6) into one net. I did this in a incremental way. The preparation times, before the first epoch starts, got longer and longer. Scaling up neural nets would result in preparations times of hours instead of minutes. </p>",
  "messages": [
    {
      "id": 891077,
      "postDate": "2020-06-17T21:57:05.750Z",
      "content": "<p>The MXU is a specialized processing unit of the TPUs. They are used to perform very fast matrix operations like multiplications, which is the main computational work when training neural nets. The TPUs also have a more generalized CPU like processing unit, used to fetch and prepare data, communicate with the host, and coordinating the processing \"workflow\".</p>\n\n<p>The documentation says you should use \n- big batch sizes\n- multiple tfrec files\n- turn of the deterministic order when reading the files. (experimental_deterministic = False)\n- tf.dataset.cache()</p>\n\n<p>Most notebooks use above points, but I have not found one, which gets the MXU over 70-80% usage. </p>\n\n<p>What must be done to get the most out of the mxu or whole TPU board? Huge/deep nets? very large input/image sizes? I mean, it seams pretty impossible to saturate the MXUs of a single TPU board. What do the big companies do when learing with complete TPU-Pods?</p>\n\n<p>While digging around this, I created following notebook (<a href=\"https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once\">Link</a>), which combines seven EfficientNets (B0-B6) into one net. I did this in a incremental way. The preparation times, before the first epoch starts, got longer and longer. Scaling up neural nets would result in preparations times of hours instead of minutes. </p>",
      "rawMarkdown": "The MXU is a specialized processing unit of the TPUs. They are used to perform very fast matrix operations like multiplications, which is the main computational work when training neural nets. The TPUs also have a more generalized CPU like processing unit, used to fetch and prepare data, communicate with the host, and coordinating the processing \"workflow\".\n\nThe documentation says you should use \n- big batch sizes\n- multiple tfrec files\n- turn of the deterministic order when reading the files. (experimental_deterministic = False)\n- tf.dataset.cache()\n\nMost notebooks use above points, but I have not found one, which gets the MXU over 70-80% usage. \n\nWhat must be done to get the most out of the mxu or whole TPU board? Huge/deep nets? very large input/image sizes? I mean, it seams pretty impossible to saturate the MXUs of a single TPU board. What do the big companies do when learing with complete TPU-Pods?\n\nWhile digging around this, I created following notebook ([Link](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once)), which combines seven EfficientNets (B0-B6) into one net. I did this in a incremental way. The preparation times, before the first epoch starts, got longer and longer. Scaling up neural nets would result in preparations times of hours instead of minutes. \n",
      "votes": 5
    },
    {
      "id": 932183,
      "postDate": "2020-07-16T19:35:48.203Z",
      "content": "<p>I think it is mostly related to the model size and data input, I remember that at the Jigsaw competition using <code>XLM-RoBERTa large</code> with large batches and long sequence sizes, I had a more efficient usage of the TPUs. But there I was reading data from memory, not from <code>tfrecs</code> so that probably contributed as well.</p>",
      "rawMarkdown": "I think it is mostly related to the model size and data input, I remember that at the Jigsaw competition using `XLM-RoBERTa large` with large batches and long sequence sizes, I had a more efficient usage of the TPUs. But there I was reading data from memory, not from `tfrecs` so that probably contributed as well.",
      "votes": 2
    },
    {
      "id": 931978,
      "postDate": "2020-07-16T15:47:30.830Z",
      "content": "<p>AgentAuers, your notebook is awesome. Thanks for sharing.</p>\n\n<p>In addition to MXU, there is also idle time and we want idle time to be 0%. In Flower comp, Martin posted some experiments <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">here</a>.</p>",
      "rawMarkdown": "AgentAuers, your notebook is awesome. Thanks for sharing.\n\nIn addition to MXU, there is also idle time and we want idle time to be 0%. In Flower comp, Martin posted some experiments [here][1].\n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443",
      "votes": 2,
      "replies": [
        {
          "id": 932608,
          "postDate": "2020-07-17T07:13:15.723Z",
          "content": "<p>Thank you for the link. very usefull :-)</p>",
          "rawMarkdown": "Thank you for the link. very usefull :-)",
          "votes": 1
        }
      ]
    },
    {
      "id": 931973,
      "postDate": "2020-07-16T15:38:56.897Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 932183,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-07-16T19:35:48.203000",
      "content": "<p>I think it is mostly related to the model size and data input, I remember that at the Jigsaw competition using <code>XLM-RoBERTa large</code> with large batches and long sequence sizes, I had a more efficient usage of the TPUs. But there I was reading data from memory, not from <code>tfrecs</code> so that probably contributed as well.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 931978,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-16T15:47:30.830000",
      "content": "<p>AgentAuers, your notebook is awesome. Thanks for sharing.</p>\n\n<p>In addition to MXU, there is also idle time and we want idle time to be 0%. In Flower comp, Martin posted some experiments <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">here</a>.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 932608,
          "author_name": "AgentAuers",
          "author_url": "",
          "post_date": "2020-07-17T07:13:15.723000",
          "content": "<p>Thank you for the link. very usefull :-)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 931973,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-16T15:38:56.897000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "891077": "The MXU is a specialized processing unit of the TPUs. They are used to perform very fast matrix operations like multiplications, which is the main computational work when training neural nets. The TPUs also have a more generalized CPU like processing unit, used to fetch and prepare data, communicate with the host, and coordinating the processing \"workflow\".\n\nThe documentation says you should use \n- big batch sizes\n- multiple tfrec files\n- turn of the deterministic order when reading the files. (experimental_deterministic = False)\n- tf.dataset.cache()\n\nMost notebooks use above points, but I have not found one, which gets the MXU over 70-80% usage. \n\nWhat must be done to get the most out of the mxu or whole TPU board? Huge/deep nets? very large input/image sizes? I mean, it seams pretty impossible to saturate the MXUs of a single TPU board. What do the big companies do when learing with complete TPU-Pods?\n\nWhile digging around this, I created following notebook ([Link](https://www.kaggle.com/agentauers/incredible-tpus-finetune-effnetb0-b6-at-once)), which combines seven EfficientNets (B0-B6) into one net. I did this in a incremental way. The preparation times, before the first epoch starts, got longer and longer. Scaling up neural nets would result in preparations times of hours instead of minutes. \n",
    "932183": "I think it is mostly related to the model size and data input, I remember that at the Jigsaw competition using `XLM-RoBERTa large` with large batches and long sequence sizes, I had a more efficient usage of the TPUs. But there I was reading data from memory, not from `tfrecs` so that probably contributed as well.",
    "931978": "AgentAuers, your notebook is awesome. Thanks for sharing.\n\nIn addition to MXU, there is also idle time and we want idle time to be 0%. In Flower comp, Martin posted some experiments [here][1].\n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443",
    "931973": ""
  }
}