{
  "id": 187708,
  "title": "Public #87 --> Private #219 Shaken but not stirred",
  "url": "/competitions/landmark-recognition-2020/discussion/187708",
  "author_name": "",
  "post_date": "2020-09-30T02:23:40.058266400Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congrats to all Winners and participants too ! Since I joined Kaggle 1 yr ago, this is probably the most challenging competition I have joined. Some called it the \"Humpback Whale ID challenge on steroids\"</p>\n<p>You may wonder why I post this while the winners have not ? Probably they need a well deserved rest now. Each commit run takes me 6 to 9 hrs, some as much as 12 hrs. Though I would love to get a medal, I have also learned alot here, like RANSAC &amp; geometric verification of local keypoints. These lessons would serve me well in future competitions</p>\n<p>I used the base kernel as starting point. <br>\nWhat I tried mostly worked in debug mode with small sample size, but not full run (not sure why) :<br>\n1)  tweaked settings in pydegensac<br>\n2) use Rapids cuML &amp; cuDF --&gt; observed very good speedup in debug mode<br>\n3) tried \"winning strategies\" from past year winners  (eg filter scores lower than Top 20k as non-landmarks)</p>\n<p>What puzzles me:<br>\n1) Highly random nature of pydegensac --&gt; some near identical runs can differ as much as 0.0050 in LB score<br>\n2) kernel timeout --&gt; I got timeout for 2 continuous weeks (Ran same code which previously worked, but not now). I had to fork another kernel &amp; restart again with 2 weeks left<br>\n3) Lack of feedback on full run --&gt; Things mostly worked when I run on debug mode, with sample size 100 to 1000. But when I launched full run, I encountered timeout, Kaggle error, csv error w/o fully understanding what happened<br>\n4) Public &amp; Private LB don't correlate --&gt; some of my \"medal winning\" kernels have quite low public LB, which I will never select</p>",
  "messages": [
    {
      "id": "1032158",
      "postDate": "09/30/2020 02:23:40",
      "content": "<p>Congrats to all Winners and participants too ! Since I joined Kaggle 1 yr ago, this is probably the most challenging competition I have joined. Some called it the \"Humpback Whale ID challenge on steroids\"</p>\n<p>You may wonder why I post this while the winners have not ? Probably they need a well deserved rest now. Each commit run takes me 6 to 9 hrs, some as much as 12 hrs. Though I would love to get a medal, I have also learned alot here, like RANSAC &amp; geometric verification of local keypoints. These lessons would serve me well in future competitions</p>\n<p>I used the base kernel as starting point. <br>\nWhat I tried mostly worked in debug mode with small sample size, but not full run (not sure why) :<br>\n1)  tweaked settings in pydegensac<br>\n2) use Rapids cuML &amp; cuDF --&gt; observed very good speedup in debug mode<br>\n3) tried \"winning strategies\" from past year winners  (eg filter scores lower than Top 20k as non-landmarks)</p>\n<p>What puzzles me:<br>\n1) Highly random nature of pydegensac --&gt; some near identical runs can differ as much as 0.0050 in LB score<br>\n2) kernel timeout --&gt; I got timeout for 2 continuous weeks (Ran same code which previously worked, but not now). I had to fork another kernel &amp; restart again with 2 weeks left<br>\n3) Lack of feedback on full run --&gt; Things mostly worked when I run on debug mode, with sample size 100 to 1000. But when I launched full run, I encountered timeout, Kaggle error, csv error w/o fully understanding what happened<br>\n4) Public &amp; Private LB don't correlate --&gt; some of my \"medal winning\" kernels have quite low public LB, which I will never select</p>",
      "rawMarkdown": "Congrats to all Winners and participants too ! Since I joined Kaggle 1 yr ago, this is probably the most challenging competition I have joined. Some called it the \"Humpback Whale ID challenge on steroids\"\n\nYou may wonder why I post this while the winners have not ? Probably they need a well deserved rest now. Each commit run takes me 6 to 9 hrs, some as much as 12 hrs. Though I would love to get a medal, I have also learned alot here, like RANSAC & geometric verification of local keypoints. These lessons would serve me well in future competitions\n\nI used the base kernel as starting point. \nWhat I tried mostly worked in debug mode with small sample size, but not full run (not sure why) :\n1)  tweaked settings in pydegensac\n2) use Rapids cuML & cuDF --> observed very good speedup in debug mode\n3) tried \"winning strategies\" from past year winners  (eg filter scores lower than Top 20k as non-landmarks)\n\nWhat puzzles me:\n1) Highly random nature of pydegensac --> some near identical runs can differ as much as 0.0050 in LB score\n2) kernel timeout --> I got timeout for 2 continuous weeks (Ran same code which previously worked, but not now). I had to fork another kernel & restart again with 2 weeks left\n3) Lack of feedback on full run --> Things mostly worked when I run on debug mode, with sample size 100 to 1000. But when I launched full run, I encountered timeout, Kaggle error, csv error w/o fully understanding what happened\n4) Public & Private LB don't correlate --> some of my \"medal winning\" kernels have quite low public LB, which I will never select",
      "votes": null
    },
    {
      "id": "1032434",
      "postDate": "09/30/2020 07:46:51",
      "content": "<p>My best kernel on the public leaderboard was not my best private score either. I think my top public scoring kernel was the result of too much RANSAC tuning/luck and that didn't help for the private test set. Having been here before, I chose the best scoring kernel that was significantly different from that top scoring kernel as my second choice for the final score. I'm glad I did as that was the better private score.</p>\n<p>Agreed, the randomness of pydegensac and the lack of feedback on submission runs did make this a particular challenge.</p>",
      "rawMarkdown": "My best kernel on the public leaderboard was not my best private score either. I think my top public scoring kernel was the result of too much RANSAC tuning/luck and that didn't help for the private test set. Having been here before, I chose the best scoring kernel that was significantly different from that top scoring kernel as my second choice for the final score. I'm glad I did as that was the better private score.\n\nAgreed, the randomness of pydegensac and the lack of feedback on submission runs did make this a particular challenge.",
      "votes": null
    },
    {
      "id": "1032453",
      "postDate": "09/30/2020 08:02:01",
      "content": "<p>Thanks for advice. My highest Public LB score is based on pyransac tuning, I had no idea it would score so low on private LB, as the shakeout was not severe in past years' Google Landmark competitions</p>\n<p>I had 2 other kernels (both not selected) --&gt; both scored Public LB 0.4879, but 1 got private LB 0.4603 while the other private LB 0.4652 (in medal zone). Maybe a lack of validation strategy cost me in this competition (I didn't know how to validate such a large dataset)</p>\n<p>cheers<br>\nsid</p>",
      "rawMarkdown": "Thanks for advice. My highest Public LB score is based on pyransac tuning, I had no idea it would score so low on private LB, as the shakeout was not severe in past years' Google Landmark competitions\n\nI had 2 other kernels (both not selected) --> both scored Public LB 0.4879, but 1 got private LB 0.4603 while the other private LB 0.4652 (in medal zone). Maybe a lack of validation strategy cost me in this competition (I didn't know how to validate such a large dataset)\n\ncheers\nsid",
      "votes": null
    },
    {
      "id": "1032594",
      "postDate": "09/30/2020 10:04:16",
      "content": "<p>Tuning RANSAC parameters is the easiest way to overfit in this competition I guess.</p>",
      "rawMarkdown": "Tuning RANSAC parameters is the easiest way to overfit in this competition I guess.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1032434,
      "author_name": "andypenrose",
      "author_url": "",
      "post_date": "09/30/2020 07:46:51",
      "content": "<p>My best kernel on the public leaderboard was not my best private score either. I think my top public scoring kernel was the result of too much RANSAC tuning/luck and that didn't help for the private test set. Having been here before, I chose the best scoring kernel that was significantly different from that top scoring kernel as my second choice for the final score. I'm glad I did as that was the better private score.</p>\n<p>Agreed, the randomness of pydegensac and the lack of feedback on submission runs did make this a particular challenge.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1032453,
          "author_name": "sidneyng",
          "author_url": "",
          "post_date": "09/30/2020 08:02:01",
          "content": "<p>Thanks for advice. My highest Public LB score is based on pyransac tuning, I had no idea it would score so low on private LB, as the shakeout was not severe in past years' Google Landmark competitions</p>\n<p>I had 2 other kernels (both not selected) --&gt; both scored Public LB 0.4879, but 1 got private LB 0.4603 while the other private LB 0.4652 (in medal zone). Maybe a lack of validation strategy cost me in this competition (I didn't know how to validate such a large dataset)</p>\n<p>cheers<br>\nsid</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032594,
          "author_name": "hav4ik",
          "author_url": "",
          "post_date": "09/30/2020 10:04:16",
          "content": "<p>Tuning RANSAC parameters is the easiest way to overfit in this competition I guess.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1032158": "Congrats to all Winners and participants too ! Since I joined Kaggle 1 yr ago, this is probably the most challenging competition I have joined. Some called it the \"Humpback Whale ID challenge on steroids\"\n\nYou may wonder why I post this while the winners have not ? Probably they need a well deserved rest now. Each commit run takes me 6 to 9 hrs, some as much as 12 hrs. Though I would love to get a medal, I have also learned alot here, like RANSAC & geometric verification of local keypoints. These lessons would serve me well in future competitions\n\nI used the base kernel as starting point. \nWhat I tried mostly worked in debug mode with small sample size, but not full run (not sure why) :\n1)  tweaked settings in pydegensac\n2) use Rapids cuML & cuDF --> observed very good speedup in debug mode\n3) tried \"winning strategies\" from past year winners  (eg filter scores lower than Top 20k as non-landmarks)\n\nWhat puzzles me:\n1) Highly random nature of pydegensac --> some near identical runs can differ as much as 0.0050 in LB score\n2) kernel timeout --> I got timeout for 2 continuous weeks (Ran same code which previously worked, but not now). I had to fork another kernel & restart again with 2 weeks left\n3) Lack of feedback on full run --> Things mostly worked when I run on debug mode, with sample size 100 to 1000. But when I launched full run, I encountered timeout, Kaggle error, csv error w/o fully understanding what happened\n4) Public & Private LB don't correlate --> some of my \"medal winning\" kernels have quite low public LB, which I will never select",
    "1032434": "My best kernel on the public leaderboard was not my best private score either. I think my top public scoring kernel was the result of too much RANSAC tuning/luck and that didn't help for the private test set. Having been here before, I chose the best scoring kernel that was significantly different from that top scoring kernel as my second choice for the final score. I'm glad I did as that was the better private score.\n\nAgreed, the randomness of pydegensac and the lack of feedback on submission runs did make this a particular challenge.",
    "1032453": "Thanks for advice. My highest Public LB score is based on pyransac tuning, I had no idea it would score so low on private LB, as the shakeout was not severe in past years' Google Landmark competitions\n\nI had 2 other kernels (both not selected) --> both scored Public LB 0.4879, but 1 got private LB 0.4603 while the other private LB 0.4652 (in medal zone). Maybe a lack of validation strategy cost me in this competition (I didn't know how to validate such a large dataset)\n\ncheers\nsid",
    "1032594": "Tuning RANSAC parameters is the easiest way to overfit in this competition I guess."
  },
  "source": "meta"
}