{
  "id": 82380,
  "title": "Approximation solution for LAP, 30th place solution",
  "url": "/competitions/humpback-whale-identification/writeups/lola-approximation-solution-for-lap-30th-place-sol",
  "author_name": "",
  "post_date": "2019-03-01T03:40:04.244816900Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Many thanks to the organizers and Kaggle for this very interesting competition, and congratulations to all winners. And thanks to <a href=\"https://www.kaggle.com/martinpiotte\">@martinpiotte</a> for his great work, we learned so much from it. </p>\n\n<p>Our approach was based on Martin's awesome solution. One of the issues with Martin's original solution that he acknowledged was the computation time for solving the LAP problem, which is a O(n^3) problem, and is the limiting step in training time. Furthermore since we are solving LAP for a score matrix with added randomness to it, we don't really need to precisely solve it. A greedy approximation to the LAP problem would be more than suffice, as discussed in <a href=\"https://antimatroid.wordpress.com/2017/03/21/a-greedy-approximation-algorithm-for-the-linear-assignment-problem/\">this link</a>. This greedy approach is a O(n^2 log(n)) approach, but we don't even need to be this precise. We ended up using a complete random approach where we search for minimum on a random permutation of rows, and after each row search remove the corresponding rows/column from the matrix. This results in an O(n^2) approximation that can be calculated in a few seconds.</p>\n\n<p>With bootstrap we were able to get a single model to 0.915 public LB, and with ensemble we got to 0.944 public LB. That seems to be the limit for Martin's approach, as observed also by @interneuron in the 24th place write up. We should have diversified our model structures more. </p>\n\n<p>One thing we noticed after the private LB comes out is that ensembling through voting can overfit the public lb quite easily, whereas ensembling through score outputs results in much less separation between public and private LB.</p>",
  "messages": [
    {
      "id": "481104",
      "postDate": "03/01/2019 03:40:04",
      "content": "<p>Many thanks to the organizers and Kaggle for this very interesting competition, and congratulations to all winners. And thanks to <a href=\"https://www.kaggle.com/martinpiotte\">@martinpiotte</a> for his great work, we learned so much from it. </p>\n\n<p>Our approach was based on Martin's awesome solution. One of the issues with Martin's original solution that he acknowledged was the computation time for solving the LAP problem, which is a O(n^3) problem, and is the limiting step in training time. Furthermore since we are solving LAP for a score matrix with added randomness to it, we don't really need to precisely solve it. A greedy approximation to the LAP problem would be more than suffice, as discussed in <a href=\"https://antimatroid.wordpress.com/2017/03/21/a-greedy-approximation-algorithm-for-the-linear-assignment-problem/\">this link</a>. This greedy approach is a O(n^2 log(n)) approach, but we don't even need to be this precise. We ended up using a complete random approach where we search for minimum on a random permutation of rows, and after each row search remove the corresponding rows/column from the matrix. This results in an O(n^2) approximation that can be calculated in a few seconds.</p>\n\n<p>With bootstrap we were able to get a single model to 0.915 public LB, and with ensemble we got to 0.944 public LB. That seems to be the limit for Martin's approach, as observed also by @interneuron in the 24th place write up. We should have diversified our model structures more. </p>\n\n<p>One thing we noticed after the private LB comes out is that ensembling through voting can overfit the public lb quite easily, whereas ensembling through score outputs results in much less separation between public and private LB.</p>",
      "rawMarkdown": "Many thanks to the organizers and Kaggle for this very interesting competition, and congratulations to all winners. And thanks to [@martinpiotte][1] for his great work, we learned so much from it. \n\nOur approach was based on Martin's awesome solution. One of the issues with Martin's original solution that he acknowledged was the computation time for solving the LAP problem, which is a O(n^3) problem, and is the limiting step in training time. Furthermore since we are solving LAP for a score matrix with added randomness to it, we don't really need to precisely solve it. A greedy approximation to the LAP problem would be more than suffice, as discussed in [this link][2]. This greedy approach is a O(n^2 log(n)) approach, but we don't even need to be this precise. We ended up using a complete random approach where we search for minimum on a random permutation of rows, and after each row search remove the corresponding rows/column from the matrix. This results in an O(n^2) approximation that can be calculated in a few seconds.\n\nWith bootstrap we were able to get a single model to 0.915 public LB, and with ensemble we got to 0.944 public LB. That seems to be the limit for Martin's approach, as observed also by @interneuron in the 24th place write up. We should have diversified our model structures more. \n\nOne thing we noticed after the private LB comes out is that ensembling through voting can overfit the public lb quite easily, whereas ensembling through score outputs results in much less separation between public and private LB.\n\n\n  [1]: https://www.kaggle.com/martinpiotte\n  [2]: https://antimatroid.wordpress.com/2017/03/21/a-greedy-approximation-algorithm-for-the-linear-assignment-problem/",
      "votes": null
    },
    {
      "id": "481140",
      "postDate": "03/01/2019 04:55:32",
      "content": "<p>Really interesting, your fast lap approximation sounds quite powerful. It is nice to see others also found similar limits with the Siamese net on these data, though I suspect the label noise probably had a big part in this limit here. In any event, I think the siamese is really powerful in general, much more so than I appreciated before this competition and a fast and accurate lap solver will really expand its uses. Great work!</p>",
      "rawMarkdown": "Really interesting, your fast lap approximation sounds quite powerful. It is nice to see others also found similar limits with the Siamese net on these data, though I suspect the label noise probably had a big part in this limit here. In any event, I think the siamese is really powerful in general, much more so than I appreciated before this competition and a fast and accurate lap solver will really expand its uses. Great work!",
      "votes": null
    },
    {
      "id": "481241",
      "postDate": "03/01/2019 07:29:31",
      "content": "<p>Great work <a href=\"/mingzhao03\">@mingzhao03</a></p>",
      "rawMarkdown": "Great work @mingzhao03",
      "votes": null
    },
    {
      "id": "481337",
      "postDate": "03/01/2019 09:27:20",
      "content": "<p>brilliant approach! I preferred prototypical network for its speed but siamese was indeed the great option through this competition. Can I ask you some pseudo code for the random search if you don't mind? </p>",
      "rawMarkdown": "brilliant approach! I preferred prototypical network for its speed but siamese was indeed the great option through this competition. Can I ask you some pseudo code for the random search if you don't mind?",
      "votes": null
    },
    {
      "id": "481470",
      "postDate": "03/01/2019 13:05:00",
      "content": "<p>Congrats <a href=\"/mingzhao03\">@mingzhao03</a> and team. Thanks for sharing.</p>",
      "rawMarkdown": "Congrats @mingzhao03 and team. Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 481140,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "03/01/2019 04:55:32",
      "content": "<p>Really interesting, your fast lap approximation sounds quite powerful. It is nice to see others also found similar limits with the Siamese net on these data, though I suspect the label noise probably had a big part in this limit here. In any event, I think the siamese is really powerful in general, much more so than I appreciated before this competition and a fast and accurate lap solver will really expand its uses. Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481241,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "03/01/2019 07:29:31",
      "content": "<p>Great work <a href=\"/mingzhao03\">@mingzhao03</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481337,
      "author_name": "soonhwankwon",
      "author_url": "",
      "post_date": "03/01/2019 09:27:20",
      "content": "<p>brilliant approach! I preferred prototypical network for its speed but siamese was indeed the great option through this competition. Can I ask you some pseudo code for the random search if you don't mind? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481470,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/01/2019 13:05:00",
      "content": "<p>Congrats <a href=\"/mingzhao03\">@mingzhao03</a> and team. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "481104": "Many thanks to the organizers and Kaggle for this very interesting competition, and congratulations to all winners. And thanks to [@martinpiotte][1] for his great work, we learned so much from it. \n\nOur approach was based on Martin's awesome solution. One of the issues with Martin's original solution that he acknowledged was the computation time for solving the LAP problem, which is a O(n^3) problem, and is the limiting step in training time. Furthermore since we are solving LAP for a score matrix with added randomness to it, we don't really need to precisely solve it. A greedy approximation to the LAP problem would be more than suffice, as discussed in [this link][2]. This greedy approach is a O(n^2 log(n)) approach, but we don't even need to be this precise. We ended up using a complete random approach where we search for minimum on a random permutation of rows, and after each row search remove the corresponding rows/column from the matrix. This results in an O(n^2) approximation that can be calculated in a few seconds.\n\nWith bootstrap we were able to get a single model to 0.915 public LB, and with ensemble we got to 0.944 public LB. That seems to be the limit for Martin's approach, as observed also by @interneuron in the 24th place write up. We should have diversified our model structures more. \n\nOne thing we noticed after the private LB comes out is that ensembling through voting can overfit the public lb quite easily, whereas ensembling through score outputs results in much less separation between public and private LB.\n\n\n  [1]: https://www.kaggle.com/martinpiotte\n  [2]: https://antimatroid.wordpress.com/2017/03/21/a-greedy-approximation-algorithm-for-the-linear-assignment-problem/",
    "481140": "Really interesting, your fast lap approximation sounds quite powerful. It is nice to see others also found similar limits with the Siamese net on these data, though I suspect the label noise probably had a big part in this limit here. In any event, I think the siamese is really powerful in general, much more so than I appreciated before this competition and a fast and accurate lap solver will really expand its uses. Great work!",
    "481241": "Great work @mingzhao03",
    "481337": "brilliant approach! I preferred prototypical network for its speed but siamese was indeed the great option through this competition. Can I ask you some pseudo code for the random search if you don't mind?",
    "481470": "Congrats @mingzhao03 and team. Thanks for sharing."
  },
  "source": "meta"
}