{
  "id": 77994,
  "title": "Cuda Hungarian LAP algorithm on GPU",
  "url": "/competitions/humpback-whale-identification/discussion/77994",
  "author_name": "",
  "post_date": "2019-01-18T12:46:03.146162100Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am considering to try this implementation of LAP for GPU : <a href=\"https://github.com/paclopes/HungarianGPU\">https://github.com/paclopes/HungarianGPU</a></p>\n\n<p>If someone wants to try it as well, let's share our performance improvement over default lap-jv in python.</p>",
  "messages": [
    {
      "id": "457969",
      "postDate": "01/18/2019 12:46:03",
      "content": "<p>I am considering to try this implementation of LAP for GPU : <a href=\"https://github.com/paclopes/HungarianGPU\">https://github.com/paclopes/HungarianGPU</a></p>\n\n<p>If someone wants to try it as well, let's share our performance improvement over default lap-jv in python.</p>",
      "rawMarkdown": "I am considering to try this implementation of LAP for GPU : https://github.com/paclopes/HungarianGPU\n\nIf someone wants to try it as well, let's share our performance improvement over default lap-jv in python.",
      "votes": null
    },
    {
      "id": "459234",
      "postDate": "01/21/2019 12:24:25",
      "content": "<p>Wow, so basically .cu is the language that programming directly on CUDA right? It's cool, thanks for sharing your work!</p>",
      "rawMarkdown": "Wow, so basically .cu is the language that programming directly on CUDA right? It's cool, thanks for sharing your work!",
      "votes": null
    },
    {
      "id": "460986",
      "postDate": "01/24/2019 23:52:36",
      "content": "<p>Interesting!\nIf you get any positive or negative outcome please share it with us. The method to solve this competition is extremely interesting, because there are many similar unbalanced problems.</p>\n\n<p>Currently the limitation of speed and number of samples is because of the LAP algorithm. And it would be very interesting if this cuda code works. </p>\n\n<p>However, the more interesting solution imo is to find a better alternative. LP is a well-studied topic in convex optimization theory, it is known to have weakly-polynomial solutions which typically run in \nO(n^3) time where n is the number of variables (problem’s dimensionality). There often are approximate algorithms which run in linear time. Those algorithms deal with matrix multiplications and can be parallelized efficiently. A programmer’s heaven! \nsource : \n <a href=\"https://blog.sourced.tech/post/lapjv/\">https://blog.sourced.tech/post/lapjv/</a> </p>\n\n<p>Also for problems with sample size more than the current dataset, finding alternative algorithm to this O(n^3) LAPjv algorithm become necessary for a reasonable training time. </p>\n\n<p>Martin mentioned that too:\nMore than 50% of the time is spent solving the Linear Assignment Problem because the algorithm used has complexity O(n^3) and provides an exact solution. However, the score matrix is randomized, so investing a lot of time to compute an exact solution to a randomized input is wasteful. In the context of this competition, with no constraint on runtime and a small dataset, it was a pragmatic choice. However, to scale this approach, a less costly randomized matching heuristic would be more effective.</p>",
      "rawMarkdown": "Interesting!\nIf you get any positive or negative outcome please share it with us. The method to solve this competition is extremely interesting, because there are many similar unbalanced problems.\n\nCurrently the limitation of speed and number of samples is because of the LAP algorithm. And it would be very interesting if this cuda code works. \n\nHowever, the more interesting solution imo is to find a better alternative. LP is a well-studied topic in convex optimization theory, it is known to have weakly-polynomial solutions which typically run in \nO(n^3) time where n is the number of variables (problem’s dimensionality). There often are approximate algorithms which run in linear time. Those algorithms deal with matrix multiplications and can be parallelized efficiently. A programmer’s heaven! \nsource : \n https://blog.sourced.tech/post/lapjv/ \n\nAlso for problems with sample size more than the current dataset, finding alternative algorithm to this O(n^3) LAPjv algorithm become necessary for a reasonable training time. \n\nMartin mentioned that too:\nMore than 50% of the time is spent solving the Linear Assignment Problem because the algorithm used has complexity O(n^3) and provides an exact solution. However, the score matrix is randomized, so investing a lot of time to compute an exact solution to a randomized input is wasteful. In the context of this competition, with no constraint on runtime and a small dataset, it was a pragmatic choice. However, to scale this approach, a less costly randomized matching heuristic would be more effective.",
      "votes": null
    },
    {
      "id": "463352",
      "postDate": "01/29/2019 21:02:59",
      "content": "<p>It is actually slower than CPU implementations of lapjv. At least on my modest hardware (GTX960). For a test matrix of size ~(14Kx14K) it worked ~3 times slower than lapjv. Maybe some high-end GPU will work faster, but not orders of magnitude faster. This implementation works on matrices with dimensions equal to the closest power of 2. So, for matrices &gt;8196 &lt; 16384 it always extends your matrix to dim = 16384. Also judging by the source code it was not even tested on matrices &gt; 8196. In the original paper authors mention that their implementation could be parallelized to several GPUs, but it seems like it isn't worth it even in multi-GPU setup. Maybe someone had different experience?</p>",
      "rawMarkdown": "It is actually slower than CPU implementations of lapjv. At least on my modest hardware (GTX960). For a test matrix of size ~(14Kx14K) it worked ~3 times slower than lapjv. Maybe some high-end GPU will work faster, but not orders of magnitude faster. This implementation works on matrices with dimensions equal to the closest power of 2. So, for matrices &gt;8196 &lt; 16384 it always extends your matrix to dim = 16384. Also judging by the source code it was not even tested on matrices &gt; 8196. In the original paper authors mention that their implementation could be parallelized to several GPUs, but it seems like it isn't worth it even in multi-GPU setup. Maybe someone had different experience?",
      "votes": null
    },
    {
      "id": "465679",
      "postDate": "02/03/2019 19:06:42",
      "content": "<p>CUDA is framework that allow you to use GPU computations. It is extension of c++ language and therefore has .cu files to allow compiler determine which files should it parse as regular c++ or with CUDA additions.</p>",
      "rawMarkdown": "CUDA is framework that allow you to use GPU computations. It is extension of c++ language and therefore has .cu files to allow compiler determine which files should it parse as regular c++ or with CUDA additions.",
      "votes": null
    },
    {
      "id": "466274",
      "postDate": "02/05/2019 02:17:25",
      "content": "<p><a href=\"/mpyrozhok\">@mpyrozhok</a>\nIf you can share a benchmark code to replicate your running time benchmark on different gpu, I can check it on my 1080ti and V100.</p>\n\n<p>Sometimes different arch of gpus make significant difference. </p>",
      "rawMarkdown": "mpyrozhok\nIf you can share a benchmark code to replicate your running time benchmark on different gpu, I can check it on my 1080ti and V100.\n\nSometimes different arch of gpus make significant difference.",
      "votes": null
    },
    {
      "id": "701950",
      "postDate": "12/24/2019 05:08:30",
      "content": "<p>I checked for upto ~8kx8k and it was working faster than Python lapjv.\nTime taken:\nHungarian GPU : 200 ms\nPython lapjv : 2 seconds\nHowever for me, the cuda code throws an error if I input user_n more than 8192. Any ideas on how to fix the error?\nAlso, how can I work with an input matrix as opposed to randomly generating the values?</p>",
      "rawMarkdown": "I checked for upto ~8kx8k and it was working faster than Python lapjv.\nTime taken:\nHungarian GPU : 200 ms\nPython lapjv : 2 seconds\nHowever for me, the cuda code throws an error if I input user_n more than 8192. Any ideas on how to fix the error?\nAlso, how can I work with an input matrix as opposed to randomly generating the values?",
      "votes": null
    },
    {
      "id": "713080",
      "postDate": "01/07/2020 21:23:30",
      "content": "<p><a href=\"/anindya42\">@anindya42</a>, The actual matrices that required processing in this competition took ~10-15 minutes to process using CPU and even more using Kaggle's P100 GPU.\nThe fix for matrices &gt; 8196 is fairly trivial and involves fixing bit-shift calculations somewhere at the start of the main file IIRC.\nSpeaking about arbitrary input matrix, I've changed original code in a way that allowed loading input data from a file in a binary format. Unfortunately I can't provide my code because tests did not show any advantage of this approach and I've lost that code somewhere along the way.</p>",
      "rawMarkdown": "anindya42, The actual matrices that required processing in this competition took ~10-15 minutes to process using CPU and even more using Kaggle's P100 GPU.\nThe fix for matrices &gt; 8196 is fairly trivial and involves fixing bit-shift calculations somewhere at the start of the main file IIRC.\nSpeaking about arbitrary input matrix, I've changed original code in a way that allowed loading input data from a file in a binary format. Unfortunately I can't provide my code because tests did not show any advantage of this approach and I've lost that code somewhere along the way.",
      "votes": null
    },
    {
      "id": "742815",
      "postDate": "02/11/2020 15:00:38",
      "content": "<p><a href=\"/anindya42\">@anindya42</a> could you provide the code you wrote which invokes the compiled output of \"HungarianCUDA.cu\"?</p>",
      "rawMarkdown": "anindya42 could you provide the code you wrote which invokes the compiled output of \"HungarianCUDA.cu\"?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 459234,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "01/21/2019 12:24:25",
      "content": "<p>Wow, so basically .cu is the language that programming directly on CUDA right? It's cool, thanks for sharing your work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 465679,
          "author_name": "vlad0922",
          "author_url": "",
          "post_date": "02/03/2019 19:06:42",
          "content": "<p>CUDA is framework that allow you to use GPU computations. It is extension of c++ language and therefore has .cu files to allow compiler determine which files should it parse as regular c++ or with CUDA additions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 460986,
      "author_name": "hwasiti",
      "author_url": "",
      "post_date": "01/24/2019 23:52:36",
      "content": "<p>Interesting!\nIf you get any positive or negative outcome please share it with us. The method to solve this competition is extremely interesting, because there are many similar unbalanced problems.</p>\n\n<p>Currently the limitation of speed and number of samples is because of the LAP algorithm. And it would be very interesting if this cuda code works. </p>\n\n<p>However, the more interesting solution imo is to find a better alternative. LP is a well-studied topic in convex optimization theory, it is known to have weakly-polynomial solutions which typically run in \nO(n^3) time where n is the number of variables (problem’s dimensionality). There often are approximate algorithms which run in linear time. Those algorithms deal with matrix multiplications and can be parallelized efficiently. A programmer’s heaven! \nsource : \n <a href=\"https://blog.sourced.tech/post/lapjv/\">https://blog.sourced.tech/post/lapjv/</a> </p>\n\n<p>Also for problems with sample size more than the current dataset, finding alternative algorithm to this O(n^3) LAPjv algorithm become necessary for a reasonable training time. </p>\n\n<p>Martin mentioned that too:\nMore than 50% of the time is spent solving the Linear Assignment Problem because the algorithm used has complexity O(n^3) and provides an exact solution. However, the score matrix is randomized, so investing a lot of time to compute an exact solution to a randomized input is wasteful. In the context of this competition, with no constraint on runtime and a small dataset, it was a pragmatic choice. However, to scale this approach, a less costly randomized matching heuristic would be more effective.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463352,
      "author_name": "mpyrozhok",
      "author_url": "",
      "post_date": "01/29/2019 21:02:59",
      "content": "<p>It is actually slower than CPU implementations of lapjv. At least on my modest hardware (GTX960). For a test matrix of size ~(14Kx14K) it worked ~3 times slower than lapjv. Maybe some high-end GPU will work faster, but not orders of magnitude faster. This implementation works on matrices with dimensions equal to the closest power of 2. So, for matrices &gt;8196 &lt; 16384 it always extends your matrix to dim = 16384. Also judging by the source code it was not even tested on matrices &gt; 8196. In the original paper authors mention that their implementation could be parallelized to several GPUs, but it seems like it isn't worth it even in multi-GPU setup. Maybe someone had different experience?</p>",
      "votes": null,
      "replies": [
        {
          "id": 701950,
          "author_name": "anindya42",
          "author_url": "",
          "post_date": "12/24/2019 05:08:30",
          "content": "<p>I checked for upto ~8kx8k and it was working faster than Python lapjv.\nTime taken:\nHungarian GPU : 200 ms\nPython lapjv : 2 seconds\nHowever for me, the cuda code throws an error if I input user_n more than 8192. Any ideas on how to fix the error?\nAlso, how can I work with an input matrix as opposed to randomly generating the values?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713080,
          "author_name": "mpyrozhok",
          "author_url": "",
          "post_date": "01/07/2020 21:23:30",
          "content": "<p><a href=\"/anindya42\">@anindya42</a>, The actual matrices that required processing in this competition took ~10-15 minutes to process using CPU and even more using Kaggle's P100 GPU.\nThe fix for matrices &gt; 8196 is fairly trivial and involves fixing bit-shift calculations somewhere at the start of the main file IIRC.\nSpeaking about arbitrary input matrix, I've changed original code in a way that allowed loading input data from a file in a binary format. Unfortunately I can't provide my code because tests did not show any advantage of this approach and I've lost that code somewhere along the way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742815,
          "author_name": "horace89",
          "author_url": "",
          "post_date": "02/11/2020 15:00:38",
          "content": "<p><a href=\"/anindya42\">@anindya42</a> could you provide the code you wrote which invokes the compiled output of \"HungarianCUDA.cu\"?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 466274,
      "author_name": "hwasiti",
      "author_url": "",
      "post_date": "02/05/2019 02:17:25",
      "content": "<p><a href=\"/mpyrozhok\">@mpyrozhok</a>\nIf you can share a benchmark code to replicate your running time benchmark on different gpu, I can check it on my 1080ti and V100.</p>\n\n<p>Sometimes different arch of gpus make significant difference. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "457969": "I am considering to try this implementation of LAP for GPU : https://github.com/paclopes/HungarianGPU\n\nIf someone wants to try it as well, let's share our performance improvement over default lap-jv in python.",
    "459234": "Wow, so basically .cu is the language that programming directly on CUDA right? It's cool, thanks for sharing your work!",
    "460986": "Interesting!\nIf you get any positive or negative outcome please share it with us. The method to solve this competition is extremely interesting, because there are many similar unbalanced problems.\n\nCurrently the limitation of speed and number of samples is because of the LAP algorithm. And it would be very interesting if this cuda code works. \n\nHowever, the more interesting solution imo is to find a better alternative. LP is a well-studied topic in convex optimization theory, it is known to have weakly-polynomial solutions which typically run in \nO(n^3) time where n is the number of variables (problem’s dimensionality). There often are approximate algorithms which run in linear time. Those algorithms deal with matrix multiplications and can be parallelized efficiently. A programmer’s heaven! \nsource : \n https://blog.sourced.tech/post/lapjv/ \n\nAlso for problems with sample size more than the current dataset, finding alternative algorithm to this O(n^3) LAPjv algorithm become necessary for a reasonable training time. \n\nMartin mentioned that too:\nMore than 50% of the time is spent solving the Linear Assignment Problem because the algorithm used has complexity O(n^3) and provides an exact solution. However, the score matrix is randomized, so investing a lot of time to compute an exact solution to a randomized input is wasteful. In the context of this competition, with no constraint on runtime and a small dataset, it was a pragmatic choice. However, to scale this approach, a less costly randomized matching heuristic would be more effective.",
    "463352": "It is actually slower than CPU implementations of lapjv. At least on my modest hardware (GTX960). For a test matrix of size ~(14Kx14K) it worked ~3 times slower than lapjv. Maybe some high-end GPU will work faster, but not orders of magnitude faster. This implementation works on matrices with dimensions equal to the closest power of 2. So, for matrices &gt;8196 &lt; 16384 it always extends your matrix to dim = 16384. Also judging by the source code it was not even tested on matrices &gt; 8196. In the original paper authors mention that their implementation could be parallelized to several GPUs, but it seems like it isn't worth it even in multi-GPU setup. Maybe someone had different experience?",
    "465679": "CUDA is framework that allow you to use GPU computations. It is extension of c++ language and therefore has .cu files to allow compiler determine which files should it parse as regular c++ or with CUDA additions.",
    "466274": "mpyrozhok\nIf you can share a benchmark code to replicate your running time benchmark on different gpu, I can check it on my 1080ti and V100.\n\nSometimes different arch of gpus make significant difference.",
    "701950": "I checked for upto ~8kx8k and it was working faster than Python lapjv.\nTime taken:\nHungarian GPU : 200 ms\nPython lapjv : 2 seconds\nHowever for me, the cuda code throws an error if I input user_n more than 8192. Any ideas on how to fix the error?\nAlso, how can I work with an input matrix as opposed to randomly generating the values?",
    "713080": "anindya42, The actual matrices that required processing in this competition took ~10-15 minutes to process using CPU and even more using Kaggle's P100 GPU.\nThe fix for matrices &gt; 8196 is fairly trivial and involves fixing bit-shift calculations somewhere at the start of the main file IIRC.\nSpeaking about arbitrary input matrix, I've changed original code in a way that allowed loading input data from a file in a binary format. Unfortunately I can't provide my code because tests did not show any advantage of this approach and I've lost that code somewhere along the way.",
    "742815": "anindya42 could you provide the code you wrote which invokes the compiled output of \"HungarianCUDA.cu\"?"
  },
  "source": "meta"
}