{
  "id": 246322,
  "title": "About those A100s ...",
  "url": "/competitions/bms-molecular-translation/discussion/246322",
  "author_name": "",
  "post_date": "2021-06-15T00:22:23.560060300Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I've been meaning to clarify this issue ever since we changed our team's name. TL;DR: we <strong>DID NOT</strong> use 1024 A100s in this competition. </p>\n<p>Very early into the competition it became clear to everyone that this problem would require substantial computing resources to tackle. Many teams changed their names to reflect that fact, and/or to express their aspiration regarding GPU resources they desired. In a way, we did the same - I changed our name to indicate what computing resources I'd like to use, but also to tease and play mind games with other competitors. 😉 The matter of fact is, A100s are still hard to come by, even for us who work at Nvidia. We at the KGMON team are still fortunate to have <strong>LOTS</strong> of compute at our disposal, much more than I ever dreamed about. Each one of us has access to 1-4 (depending on the demand) DGX-1 servers with 8 V100s in our internal cloud, plus a direct access to a DGX Station workstation with 4 V100s. </p>\n<p>Not too long after the nucleus of our team had formed, I was approached by the <a href=\"https://www.nvidia.com/en-us/data-center/dgx-station-a100/\" target=\"_blank\">DGX Station A100</a> marketing team to see if we could do some more testing of those wonderful machines. I offered to try one of them out for this competition, and was fortunately given access to it. </p>\n<p>DGX Station A100 is an incredible machine. It boasts 4 A100 GPUs in a single workstation tower configuration. Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car! </p>\n<p>And the DGX Station A100 for the most part delivered. Unfortunately, <a href=\"https://twitter.com/arankomatsuzaki/status/1401635261587488769\" target=\"_blank\">there is a bug in PyTorch that prevents A100s from running as fast as possible</a>, but we only found out about it after the competition ended, so were unable to push the A100s to their speed limit. Nonetheless, being able to squeeze big batch sizes into 80 GB of GPU RAM per card was definitely a big game changer. Running the same model on 4 A100 compared to 8 V100s would enable us to squeeze an extra 0.1 of performance, which in this competition ended up being crucial. </p>\n<p>Towards the end of the competition we were able to use two DGX Station A100s. That was a big help, and our best models were all trained on these machines. We had started training even bigger models, but unfortunately we had ran out of time. There is still a lot of room for improvement with these machines, and I am looking forward to the opportunities to do some advanced data science and machine learning with them in the near future. </p>\n<p><img src=\"https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-station-a100/nvidia-dgx-station-og.jpg\" alt=\"DGX Station A100\"></p>",
  "messages": [
    {
      "id": "1349605",
      "postDate": "06/15/2021 00:22:23",
      "content": "<p>I've been meaning to clarify this issue ever since we changed our team's name. TL;DR: we <strong>DID NOT</strong> use 1024 A100s in this competition. </p>\n<p>Very early into the competition it became clear to everyone that this problem would require substantial computing resources to tackle. Many teams changed their names to reflect that fact, and/or to express their aspiration regarding GPU resources they desired. In a way, we did the same - I changed our name to indicate what computing resources I'd like to use, but also to tease and play mind games with other competitors. 😉 The matter of fact is, A100s are still hard to come by, even for us who work at Nvidia. We at the KGMON team are still fortunate to have <strong>LOTS</strong> of compute at our disposal, much more than I ever dreamed about. Each one of us has access to 1-4 (depending on the demand) DGX-1 servers with 8 V100s in our internal cloud, plus a direct access to a DGX Station workstation with 4 V100s. </p>\n<p>Not too long after the nucleus of our team had formed, I was approached by the <a href=\"https://www.nvidia.com/en-us/data-center/dgx-station-a100/\" target=\"_blank\">DGX Station A100</a> marketing team to see if we could do some more testing of those wonderful machines. I offered to try one of them out for this competition, and was fortunately given access to it. </p>\n<p>DGX Station A100 is an incredible machine. It boasts 4 A100 GPUs in a single workstation tower configuration. Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car! </p>\n<p>And the DGX Station A100 for the most part delivered. Unfortunately, <a href=\"https://twitter.com/arankomatsuzaki/status/1401635261587488769\" target=\"_blank\">there is a bug in PyTorch that prevents A100s from running as fast as possible</a>, but we only found out about it after the competition ended, so were unable to push the A100s to their speed limit. Nonetheless, being able to squeeze big batch sizes into 80 GB of GPU RAM per card was definitely a big game changer. Running the same model on 4 A100 compared to 8 V100s would enable us to squeeze an extra 0.1 of performance, which in this competition ended up being crucial. </p>\n<p>Towards the end of the competition we were able to use two DGX Station A100s. That was a big help, and our best models were all trained on these machines. We had started training even bigger models, but unfortunately we had ran out of time. There is still a lot of room for improvement with these machines, and I am looking forward to the opportunities to do some advanced data science and machine learning with them in the near future. </p>\n<p><img src=\"https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-station-a100/nvidia-dgx-station-og.jpg\" alt=\"DGX Station A100\"></p>",
      "rawMarkdown": "I've been meaning to clarify this issue ever since we changed our team's name. TL;DR: we **DID NOT** use 1024 A100s in this competition. \n\nVery early into the competition it became clear to everyone that this problem would require substantial computing resources to tackle. Many teams changed their names to reflect that fact, and/or to express their aspiration regarding GPU resources they desired. In a way, we did the same - I changed our name to indicate what computing resources I'd like to use, but also to tease and play mind games with other competitors. 😉 The matter of fact is, A100s are still hard to come by, even for us who work at Nvidia. We at the KGMON team are still fortunate to have **LOTS** of compute at our disposal, much more than I ever dreamed about. Each one of us has access to 1-4 (depending on the demand) DGX-1 servers with 8 V100s in our internal cloud, plus a direct access to a DGX Station workstation with 4 V100s. \n\nNot too long after the nucleus of our team had formed, I was approached by the [DGX Station A100](https://www.nvidia.com/en-us/data-center/dgx-station-a100/) marketing team to see if we could do some more testing of those wonderful machines. I offered to try one of them out for this competition, and was fortunately given access to it. \n\nDGX Station A100 is an incredible machine. It boasts 4 A100 GPUs in a single workstation tower configuration. Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car! \n\nAnd the DGX Station A100 for the most part delivered. Unfortunately, [there is a bug in PyTorch that prevents A100s from running as fast as possible](https://twitter.com/arankomatsuzaki/status/1401635261587488769), but we only found out about it after the competition ended, so were unable to push the A100s to their speed limit. Nonetheless, being able to squeeze big batch sizes into 80 GB of GPU RAM per card was definitely a big game changer. Running the same model on 4 A100 compared to 8 V100s would enable us to squeeze an extra 0.1 of performance, which in this competition ended up being crucial. \n\nTowards the end of the competition we were able to use two DGX Station A100s. That was a big help, and our best models were all trained on these machines. We had started training even bigger models, but unfortunately we had ran out of time. There is still a lot of room for improvement with these machines, and I am looking forward to the opportunities to do some advanced data science and machine learning with them in the near future. \n\n![DGX Station A100](https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-station-a100/nvidia-dgx-station-og.jpg)",
      "votes": null
    },
    {
      "id": "1349847",
      "postDate": "06/15/2021 06:03:13",
      "content": "<p>we need to make a table to shows the type and number of GPUs used by the top teams</p>",
      "rawMarkdown": "we need to make a table to shows the type and number of GPUs used by the top teams",
      "votes": null
    },
    {
      "id": "1356272",
      "postDate": "06/18/2021 21:18:14",
      "content": "<blockquote>\n  <p>Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car!</p>\n</blockquote>\n<p>Meanwhile, me with one Kaggle P100 GPU be like: <strong>HEMLP ME!!</strong></p>",
      "rawMarkdown": "> Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car!\n\nMeanwhile, me with one Kaggle P100 GPU be like: **HEMLP ME!!**",
      "votes": null
    },
    {
      "id": "1357571",
      "postDate": "06/19/2021 19:36:01",
      "content": "<p>Hi there! I'm super curious… with that hardware, how long would it usually take to infer on the test set?</p>",
      "rawMarkdown": "Hi there! I'm super curious... with that hardware, how long would it usually take to infer on the test set?",
      "votes": null
    },
    {
      "id": "1357638",
      "postDate": "06/19/2021 20:49:32",
      "content": "<p>We actually never ended up using A100s for the full test inference. It would take anywhere between 4 and 12 hours to infer on the full test set with V100s, depending on the size of the model and resolution.</p>",
      "rawMarkdown": "We actually never ended up using A100s for the full test inference. It would take anywhere between 4 and 12 hours to infer on the full test set with V100s, depending on the size of the model and resolution.",
      "votes": null
    },
    {
      "id": "1358516",
      "postDate": "06/20/2021 14:31:27",
      "content": "<p>HAHA, I also changed my team name to 'Two 2080Tis are sleeping well' inspired by 8 V100 team.<br>\nThey were literally sleeping while training because I used only kaggle and colab TPU!</p>",
      "rawMarkdown": "HAHA, I also changed my team name to 'Two 2080Tis are sleeping well' inspired by 8 V100 team.\nThey were literally sleeping while training because I used only kaggle and colab TPU!",
      "votes": null
    },
    {
      "id": "1358589",
      "postDate": "06/20/2021 15:43:33",
      "content": "<p>Haha, that’s awesome - seems like whatever you did it worked well for you.</p>",
      "rawMarkdown": "Haha, that’s awesome - seems like whatever you did it worked well for you.",
      "votes": null
    },
    {
      "id": "1365919",
      "postDate": "06/26/2021 08:49:39",
      "content": "<p>I’m disappointed that no one took Pentium 486 or Intel 8088 as team names</p>",
      "rawMarkdown": "I’m disappointed that no one took Pentium 486 or Intel 8088 as team names",
      "votes": null
    },
    {
      "id": "1376005",
      "postDate": "07/04/2021 16:41:42",
      "content": "<p>Wow, that's great <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>. You used only Kaggle and Colab TPUs<br>\nWe used almost 16+ GPUs to train + Kaggle GPU. And it took about 1 month to get to a reasonable result</p>",
      "rawMarkdown": "Wow, that's great @bamps53. You used only Kaggle and Colab TPUs\nWe used almost 16+ GPUs to train + Kaggle GPU. And it took about 1 month to get to a reasonable result",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1349847,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/15/2021 06:03:13",
      "content": "<p>we need to make a table to shows the type and number of GPUs used by the top teams</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1356272,
      "author_name": "tahsin",
      "author_url": "",
      "post_date": "06/18/2021 21:18:14",
      "content": "<blockquote>\n  <p>Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car!</p>\n</blockquote>\n<p>Meanwhile, me with one Kaggle P100 GPU be like: <strong>HEMLP ME!!</strong></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1357571,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "06/19/2021 19:36:01",
      "content": "<p>Hi there! I'm super curious… with that hardware, how long would it usually take to infer on the test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1357638,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "06/19/2021 20:49:32",
          "content": "<p>We actually never ended up using A100s for the full test inference. It would take anywhere between 4 and 12 hours to infer on the full test set with V100s, depending on the size of the model and resolution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1358516,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "06/20/2021 14:31:27",
      "content": "<p>HAHA, I also changed my team name to 'Two 2080Tis are sleeping well' inspired by 8 V100 team.<br>\nThey were literally sleeping while training because I used only kaggle and colab TPU!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1358589,
          "author_name": "tunguz",
          "author_url": "",
          "post_date": "06/20/2021 15:43:33",
          "content": "<p>Haha, that’s awesome - seems like whatever you did it worked well for you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1376005,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "07/04/2021 16:41:42",
          "content": "<p>Wow, that's great <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>. You used only Kaggle and Colab TPUs<br>\nWe used almost 16+ GPUs to train + Kaggle GPU. And it took about 1 month to get to a reasonable result</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1365919,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "06/26/2021 08:49:39",
      "content": "<p>I’m disappointed that no one took Pentium 486 or Intel 8088 as team names</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1349605": "I've been meaning to clarify this issue ever since we changed our team's name. TL;DR: we **DID NOT** use 1024 A100s in this competition. \n\nVery early into the competition it became clear to everyone that this problem would require substantial computing resources to tackle. Many teams changed their names to reflect that fact, and/or to express their aspiration regarding GPU resources they desired. In a way, we did the same - I changed our name to indicate what computing resources I'd like to use, but also to tease and play mind games with other competitors. 😉 The matter of fact is, A100s are still hard to come by, even for us who work at Nvidia. We at the KGMON team are still fortunate to have **LOTS** of compute at our disposal, much more than I ever dreamed about. Each one of us has access to 1-4 (depending on the demand) DGX-1 servers with 8 V100s in our internal cloud, plus a direct access to a DGX Station workstation with 4 V100s. \n\nNot too long after the nucleus of our team had formed, I was approached by the [DGX Station A100](https://www.nvidia.com/en-us/data-center/dgx-station-a100/) marketing team to see if we could do some more testing of those wonderful machines. I offered to try one of them out for this competition, and was fortunately given access to it. \n\nDGX Station A100 is an incredible machine. It boasts 4 A100 GPUs in a single workstation tower configuration. Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car! \n\nAnd the DGX Station A100 for the most part delivered. Unfortunately, [there is a bug in PyTorch that prevents A100s from running as fast as possible](https://twitter.com/arankomatsuzaki/status/1401635261587488769), but we only found out about it after the competition ended, so were unable to push the A100s to their speed limit. Nonetheless, being able to squeeze big batch sizes into 80 GB of GPU RAM per card was definitely a big game changer. Running the same model on 4 A100 compared to 8 V100s would enable us to squeeze an extra 0.1 of performance, which in this competition ended up being crucial. \n\nTowards the end of the competition we were able to use two DGX Station A100s. That was a big help, and our best models were all trained on these machines. We had started training even bigger models, but unfortunately we had ran out of time. There is still a lot of room for improvement with these machines, and I am looking forward to the opportunities to do some advanced data science and machine learning with them in the near future. \n\n![DGX Station A100](https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-station-a100/nvidia-dgx-station-og.jpg)",
    "1349847": "we need to make a table to shows the type and number of GPUs used by the top teams",
    "1356272": "> Having that kind of computing beast sitting under your desk is like driving around the town on a jet-engine souped up sports car!\n\nMeanwhile, me with one Kaggle P100 GPU be like: **HEMLP ME!!**",
    "1357571": "Hi there! I'm super curious... with that hardware, how long would it usually take to infer on the test set?",
    "1357638": "We actually never ended up using A100s for the full test inference. It would take anywhere between 4 and 12 hours to infer on the full test set with V100s, depending on the size of the model and resolution.",
    "1358516": "HAHA, I also changed my team name to 'Two 2080Tis are sleeping well' inspired by 8 V100 team.\nThey were literally sleeping while training because I used only kaggle and colab TPU!",
    "1358589": "Haha, that’s awesome - seems like whatever you did it worked well for you.",
    "1365919": "I’m disappointed that no one took Pentium 486 or Intel 8088 as team names",
    "1376005": "Wow, that's great @bamps53. You used only Kaggle and Colab TPUs\nWe used almost 16+ GPUs to train + Kaggle GPU. And it took about 1 month to get to a reasonable result"
  },
  "source": "meta"
}