{
  "id": 501246,
  "title": "Running inference of multiple models in parellel on OpenVino",
  "url": "/competitions/birdclef-2024/discussion/501246",
  "author_name": "Aurelio",
  "post_date": "2024-05-08T15:32:50.995000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So I've been playing around with openvino a bit to try and improve the inference time, and here are the results I got (#SETTING(Number of samples), effincientnet,resnet (inference in seconds)):</p>\n<h1>No-Hyperthreading(68) 85,155</h1>\n<h1>Hyperthreading(68) 71.81,151</h1>\n<h1>async+hyperthreading(68) 72.70,150</h1>\n<h1>Throughput hint+async+hyperthreading(68) 68.39,148</h1>\n<h1>Throughput hint+asyncQueue-1 job (68) 81.97,157</h1>\n<h1>Throughput hint+asyncQueue-4 job (68) 67,149</h1>\n<p>(Still need to try model quantization and fp16 which the cpu in theory supports)</p>\n<p>So as one can notice setting hypertreading brings a 17% speedup and the throughput hint another few percentage points.  I tried to run also difference modalities of asyncronous inference but apart from not bringing much performance improvement the predictions I got in the end where wrong (I also made sure that I was reordering the outputs in case there was some problem). Afterwards I tried many other approaches in order to run parallel inference between the two models but I wasn't that lucky. I tried different libraries like multiprocessing, joblib (from last years 2nd solution <a href=\"https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel\" target=\"_blank\">https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel</a>) and concurrent but to no avail, either I was getting a deadlock, the inference was slower or I ended up using only 2 cores and essentially running the application like it's sequential (So the CPU usage was at 200% and never at 400% like for the hypertreading case). I tried changing other global parameter like OMP_NUM_THREADS and similar ones like os.sched_setaffinity() to force the different inference process to run on separate sockets (We should have in theory 4 cores, with 2 per socket) but to no avail. Also tried to use p.cpu_affinity([worker]) from psutil but no luck there as well. </p>\n<p>Has someone managed to solve this problem and get a 400% cpu usage when running in parallel?</p>",
  "messages": [
    {
      "id": 2801353,
      "postDate": "2024-05-08T15:32:50.997Z",
      "content": "<p>So I've been playing around with openvino a bit to try and improve the inference time, and here are the results I got (#SETTING(Number of samples), effincientnet,resnet (inference in seconds)):</p>\n<h1>No-Hyperthreading(68) 85,155</h1>\n<h1>Hyperthreading(68) 71.81,151</h1>\n<h1>async+hyperthreading(68) 72.70,150</h1>\n<h1>Throughput hint+async+hyperthreading(68) 68.39,148</h1>\n<h1>Throughput hint+asyncQueue-1 job (68) 81.97,157</h1>\n<h1>Throughput hint+asyncQueue-4 job (68) 67,149</h1>\n<p>(Still need to try model quantization and fp16 which the cpu in theory supports)</p>\n<p>So as one can notice setting hypertreading brings a 17% speedup and the throughput hint another few percentage points.  I tried to run also difference modalities of asyncronous inference but apart from not bringing much performance improvement the predictions I got in the end where wrong (I also made sure that I was reordering the outputs in case there was some problem). Afterwards I tried many other approaches in order to run parallel inference between the two models but I wasn't that lucky. I tried different libraries like multiprocessing, joblib (from last years 2nd solution <a href=\"https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel\" target=\"_blank\">https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel</a>) and concurrent but to no avail, either I was getting a deadlock, the inference was slower or I ended up using only 2 cores and essentially running the application like it's sequential (So the CPU usage was at 200% and never at 400% like for the hypertreading case). I tried changing other global parameter like OMP_NUM_THREADS and similar ones like os.sched_setaffinity() to force the different inference process to run on separate sockets (We should have in theory 4 cores, with 2 per socket) but to no avail. Also tried to use p.cpu_affinity([worker]) from psutil but no luck there as well. </p>\n<p>Has someone managed to solve this problem and get a 400% cpu usage when running in parallel?</p>",
      "rawMarkdown": "So I've been playing around with openvino a bit to try and improve the inference time, and here are the results I got (#SETTING(Number of samples), effincientnet,resnet (inference in seconds)):\n\n#No-Hyperthreading(68) 85,155\n#Hyperthreading(68) 71.81,151\n#async+hyperthreading(68) 72.70,150\n#Throughput hint+async+hyperthreading(68) 68.39,148\n#Throughput hint+asyncQueue-1 job (68) 81.97,157\n#Throughput hint+asyncQueue-4 job (68) 67,149\n(Still need to try model quantization and fp16 which the cpu in theory supports)\n\nSo as one can notice setting hypertreading brings a 17% speedup and the throughput hint another few percentage points.  I tried to run also difference modalities of asyncronous inference but apart from not bringing much performance improvement the predictions I got in the end where wrong (I also made sure that I was reordering the outputs in case there was some problem). Afterwards I tried many other approaches in order to run parallel inference between the two models but I wasn't that lucky. I tried different libraries like multiprocessing, joblib (from last years 2nd solution https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel) and concurrent but to no avail, either I was getting a deadlock, the inference was slower or I ended up using only 2 cores and essentially running the application like it's sequential (So the CPU usage was at 200% and never at 400% like for the hypertreading case). I tried changing other global parameter like OMP_NUM_THREADS and similar ones like os.sched_setaffinity() to force the different inference process to run on separate sockets (We should have in theory 4 cores, with 2 per socket) but to no avail. Also tried to use p.cpu_affinity([worker]) from psutil but no luck there as well. \n\nHas someone managed to solve this problem and get a 400% cpu usage when running in parallel?",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2801353": "So I've been playing around with openvino a bit to try and improve the inference time, and here are the results I got (#SETTING(Number of samples), effincientnet,resnet (inference in seconds)):\n\n#No-Hyperthreading(68) 85,155\n#Hyperthreading(68) 71.81,151\n#async+hyperthreading(68) 72.70,150\n#Throughput hint+async+hyperthreading(68) 68.39,148\n#Throughput hint+asyncQueue-1 job (68) 81.97,157\n#Throughput hint+asyncQueue-4 job (68) 67,149\n(Still need to try model quantization and fp16 which the cpu in theory supports)\n\nSo as one can notice setting hypertreading brings a 17% speedup and the throughput hint another few percentage points.  I tried to run also difference modalities of asyncronous inference but apart from not bringing much performance improvement the predictions I got in the end where wrong (I also made sure that I was reordering the outputs in case there was some problem). Afterwards I tried many other approaches in order to run parallel inference between the two models but I wasn't that lucky. I tried different libraries like multiprocessing, joblib (from last years 2nd solution https://www.kaggle.com/code/honglihang/2nd-place-solution-inference-kernel) and concurrent but to no avail, either I was getting a deadlock, the inference was slower or I ended up using only 2 cores and essentially running the application like it's sequential (So the CPU usage was at 200% and never at 400% like for the hypertreading case). I tried changing other global parameter like OMP_NUM_THREADS and similar ones like os.sched_setaffinity() to force the different inference process to run on separate sockets (We should have in theory 4 cores, with 2 per socket) but to no avail. Also tried to use p.cpu_affinity([worker]) from psutil but no luck there as well. \n\nHas someone managed to solve this problem and get a 400% cpu usage when running in parallel?"
  }
}