{
  "id": 402493,
  "title": "Utilize OpenMP on PyTorch",
  "url": "/competitions/birdclef-2023/discussion/402493",
  "author_name": "",
  "post_date": "2023-04-18T15:42:55.584174600Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>According to the <a href=\"https://pytorch.org/tutorials/recipes/recipes/tuning_guide.html\" target=\"_blank\">PyTorch Performance Tuning Guide</a>, OpenMP is utilized to bring better performance for parallel computation tasks.</p>\n<p>In my experiments, I was able to confirm that the inference time was consistently reduced.</p>\n<h3>Setting</h3>\n<ul>\n<li>Data: <a href=\"https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-test\" target=\"_blank\">200 test_sound_scapes</a></li>\n<li>Model: eca_nfnet_l0 (JIT)</li>\n<li>Using concurrent ThreadPoolExecutor(See <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> 's nice <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/401587\" target=\"_blank\">post</a>)</li>\n</ul>\n<h3>model * 1</h3>\n<ul>\n<li>with OpenMP: 2081.7s</li>\n<li>without OpenMP: 2143.3s</li>\n</ul>\n<h3>model * 4</h3>\n<ul>\n<li>with OpenMP: 6880.0s</li>\n<li>without OpenMP: 7693.9s</li>\n</ul>\n<p>It looks like a coincidence for 1 model, but the inference times are very different when there are multiple models.</p>\n<p>Just add the code below to apply it.</p>\n<pre><code>!export OMP_NUM_THREADS=N\n\n!export OMP_SCHEDULE=STATIC\n!export OMP_PROC_BIND=CLOSE\n!export GOMP_CPU_AFFINITY=\"N-M\"\n</code></pre>\n<p>Finally, I'm no expert in this field, so please correct me if I'm wrong.</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "2226031",
      "postDate": "04/18/2023 15:42:55",
      "content": "<p>According to the <a href=\"https://pytorch.org/tutorials/recipes/recipes/tuning_guide.html\" target=\"_blank\">PyTorch Performance Tuning Guide</a>, OpenMP is utilized to bring better performance for parallel computation tasks.</p>\n<p>In my experiments, I was able to confirm that the inference time was consistently reduced.</p>\n<h3>Setting</h3>\n<ul>\n<li>Data: <a href=\"https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-test\" target=\"_blank\">200 test_sound_scapes</a></li>\n<li>Model: eca_nfnet_l0 (JIT)</li>\n<li>Using concurrent ThreadPoolExecutor(See <a href=\"https://www.kaggle.com/leonshangguan\" target=\"_blank\">@leonshangguan</a> 's nice <a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/401587\" target=\"_blank\">post</a>)</li>\n</ul>\n<h3>model * 1</h3>\n<ul>\n<li>with OpenMP: 2081.7s</li>\n<li>without OpenMP: 2143.3s</li>\n</ul>\n<h3>model * 4</h3>\n<ul>\n<li>with OpenMP: 6880.0s</li>\n<li>without OpenMP: 7693.9s</li>\n</ul>\n<p>It looks like a coincidence for 1 model, but the inference times are very different when there are multiple models.</p>\n<p>Just add the code below to apply it.</p>\n<pre><code>!export OMP_NUM_THREADS=N\n\n!export OMP_SCHEDULE=STATIC\n!export OMP_PROC_BIND=CLOSE\n!export GOMP_CPU_AFFINITY=\"N-M\"\n</code></pre>\n<p>Finally, I'm no expert in this field, so please correct me if I'm wrong.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "According to the [PyTorch Performance Tuning Guide](https://pytorch.org/tutorials/recipes/recipes/tuning_guide.html), OpenMP is utilized to bring better performance for parallel computation tasks.\n\nIn my experiments, I was able to confirm that the inference time was consistently reduced.\n\n### Setting\n- Data: [200 test_sound_scapes](https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-test)\n- Model: eca_nfnet_l0 (JIT)\n- Using concurrent ThreadPoolExecutor(See @leonshangguan 's nice [post](https://www.kaggle.com/competitions/birdclef-2023/discussion/401587))\n\n### model * 1\n- with OpenMP: 2081.7s\n- without OpenMP: 2143.3s\n\n### model * 4\n- with OpenMP: 6880.0s\n- without OpenMP: 7693.9s\n\nIt looks like a coincidence for 1 model, but the inference times are very different when there are multiple models.\n\nJust add the code below to apply it.\n\n```\n!export OMP_NUM_THREADS=N\n\n!export OMP_SCHEDULE=STATIC\n!export OMP_PROC_BIND=CLOSE\n!export GOMP_CPU_AFFINITY=\"N-M\"\n```\n\nFinally, I'm no expert in this field, so please correct me if I'm wrong.\n\nThanks!",
      "votes": null
    },
    {
      "id": "2226795",
      "postDate": "04/19/2023 08:11:16",
      "content": "<p>Were you able to set environment variables with the !export statement?<br>\nWhat are the best values of N and M?<br>\nI put the following description and the notebook that timed out finished in time.</p>\n<pre><code>import os\nos.environ[\"OMP_NUM_THREADS\"]=\"2\"\nos.environ[\"OMP_SCHEDULE\"]=\"STATIC\"\nos.environ[\"OMP_PROC_BIND\"]=\"CLOSE\"\nimport torch\ntorch.set_num_threads(4)\n</code></pre>",
      "rawMarkdown": "Were you able to set environment variables with the !export statement?\nWhat are the best values of N and M?\nI put the following description and the notebook that timed out finished in time.\n```\nimport os\nos.environ[\"OMP_NUM_THREADS\"]=\"2\"\nos.environ[\"OMP_SCHEDULE\"]=\"STATIC\"\nos.environ[\"OMP_PROC_BIND\"]=\"CLOSE\"\nimport torch\ntorch.set_num_threads(4)\n```",
      "votes": null
    },
    {
      "id": "2227134",
      "postDate": "04/19/2023 14:21:31",
      "content": "<p>Clearly you are correct. I posted this topic yesterday because I got similar results 3 times, but it won't reproduce today. I had might see an illusion😨 Thank you for correcting my mistake. And I will continue the validation with your way.</p>",
      "rawMarkdown": "Clearly you are correct. I posted this topic yesterday because I got similar results 3 times, but it won't reproduce today. I had might see an illusion😨 Thank you for correcting my mistake. And I will continue the validation with your way.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2226795,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "04/19/2023 08:11:16",
      "content": "<p>Were you able to set environment variables with the !export statement?<br>\nWhat are the best values of N and M?<br>\nI put the following description and the notebook that timed out finished in time.</p>\n<pre><code>import os\nos.environ[\"OMP_NUM_THREADS\"]=\"2\"\nos.environ[\"OMP_SCHEDULE\"]=\"STATIC\"\nos.environ[\"OMP_PROC_BIND\"]=\"CLOSE\"\nimport torch\ntorch.set_num_threads(4)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2227134,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "04/19/2023 14:21:31",
          "content": "<p>Clearly you are correct. I posted this topic yesterday because I got similar results 3 times, but it won't reproduce today. I had might see an illusion😨 Thank you for correcting my mistake. And I will continue the validation with your way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2226031": "According to the [PyTorch Performance Tuning Guide](https://pytorch.org/tutorials/recipes/recipes/tuning_guide.html), OpenMP is utilized to bring better performance for parallel computation tasks.\n\nIn my experiments, I was able to confirm that the inference time was consistently reduced.\n\n### Setting\n- Data: [200 test_sound_scapes](https://www.kaggle.com/datasets/atsunorifujita/birdclef-2023-test)\n- Model: eca_nfnet_l0 (JIT)\n- Using concurrent ThreadPoolExecutor(See @leonshangguan 's nice [post](https://www.kaggle.com/competitions/birdclef-2023/discussion/401587))\n\n### model * 1\n- with OpenMP: 2081.7s\n- without OpenMP: 2143.3s\n\n### model * 4\n- with OpenMP: 6880.0s\n- without OpenMP: 7693.9s\n\nIt looks like a coincidence for 1 model, but the inference times are very different when there are multiple models.\n\nJust add the code below to apply it.\n\n```\n!export OMP_NUM_THREADS=N\n\n!export OMP_SCHEDULE=STATIC\n!export OMP_PROC_BIND=CLOSE\n!export GOMP_CPU_AFFINITY=\"N-M\"\n```\n\nFinally, I'm no expert in this field, so please correct me if I'm wrong.\n\nThanks!",
    "2226795": "Were you able to set environment variables with the !export statement?\nWhat are the best values of N and M?\nI put the following description and the notebook that timed out finished in time.\n```\nimport os\nos.environ[\"OMP_NUM_THREADS\"]=\"2\"\nos.environ[\"OMP_SCHEDULE\"]=\"STATIC\"\nos.environ[\"OMP_PROC_BIND\"]=\"CLOSE\"\nimport torch\ntorch.set_num_threads(4)\n```",
    "2227134": "Clearly you are correct. I posted this topic yesterday because I got similar results 3 times, but it won't reproduce today. I had might see an illusion😨 Thank you for correcting my mistake. And I will continue the validation with your way."
  },
  "source": "meta"
}