{
  "id": 314418,
  "title": "confusion about improving LB scores and the model",
  "url": "/competitions/happy-whale-and-dolphin/discussion/314418",
  "author_name": "",
  "post_date": "2022-03-22T15:05:38.367474200Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Q1</h1>\n<p>I wonder <strong>how I can improve my LB scores in this competition</strong>. No need to be too specific if it's not convenient. Just some direction is enough.<br>\nMy confusion is most of the kagglers on top use the same model(as I see in public notebook), do they have special techniques <strong>except fine-tuning the parameters</strong>(or sharing some useful tricks on fine-tuning will be also grateful)?</p>\n<h1>Q2</h1>\n<p>In this competition, I see pretraining a good embedding and then use KNN is much better than directly classification.<br>\n I have a test offline, even I only classfiy  a species with 1,000 unique individuals, the MAP@5 is always no more than 0.2, obviously far more than the SOTA . Maybe it's partially because I didn't do any fine-tuning, but I wonder <strong>why the former method is much better than the latter</strong>?</p>\n<p><strong>Previously thanks for your help.</strong></p>",
  "messages": [
    {
      "id": "1731642",
      "postDate": "03/22/2022 15:05:38",
      "content": "<h1>Q1</h1>\n<p>I wonder <strong>how I can improve my LB scores in this competition</strong>. No need to be too specific if it's not convenient. Just some direction is enough.<br>\nMy confusion is most of the kagglers on top use the same model(as I see in public notebook), do they have special techniques <strong>except fine-tuning the parameters</strong>(or sharing some useful tricks on fine-tuning will be also grateful)?</p>\n<h1>Q2</h1>\n<p>In this competition, I see pretraining a good embedding and then use KNN is much better than directly classification.<br>\n I have a test offline, even I only classfiy  a species with 1,000 unique individuals, the MAP@5 is always no more than 0.2, obviously far more than the SOTA . Maybe it's partially because I didn't do any fine-tuning, but I wonder <strong>why the former method is much better than the latter</strong>?</p>\n<p><strong>Previously thanks for your help.</strong></p>",
      "rawMarkdown": "# Q1\nI wonder **how I can improve my LB scores in this competition**. No need to be too specific if it's not convenient. Just some direction is enough.\nMy confusion is most of the kagglers on top use the same model(as I see in public notebook), do they have special techniques **except fine-tuning the parameters**(or sharing some useful tricks on fine-tuning will be also grateful)?\n# Q2\nIn this competition, I see pretraining a good embedding and then use KNN is much better than directly classification.\n I have a test offline, even I only classfiy  a species with 1,000 unique individuals, the MAP@5 is always no more than 0.2, obviously far more than the SOTA . Maybe it's partially because I didn't do any fine-tuning, but I wonder **why the former method is much better than the latter**?\n\n**Previously thanks for your help.**",
      "votes": null
    },
    {
      "id": "1731685",
      "postDate": "03/22/2022 15:45:40",
      "content": "<p>You can read top solutions when the competition end.<br>\nTo improve LB there is a combination of many things: Dataset, Data augmentation, model architecture, ensembling, …</p>",
      "rawMarkdown": "You can read top solutions when the competition end.\nTo improve LB there is a combination of many things: Dataset, Data augmentation, model architecture, ensembling, ...",
      "votes": null
    },
    {
      "id": "1732058",
      "postDate": "03/23/2022 00:50:42",
      "content": "<p>But In this competition, I see many kagglers use the same Dataset, Data augmentation and model architecture, but scores various from 0.2 to 0.7 plus. Ensembling can't  imporve so much, so other improvement can all attribute to fine-tuning?</p>",
      "rawMarkdown": "But In this competition, I see many kagglers use the same Dataset, Data augmentation and model architecture, but scores various from 0.2 to 0.7 plus. Ensembling can't  imporve so much, so other improvement can all attribute to fine-tuning?",
      "votes": null
    },
    {
      "id": "1732219",
      "postDate": "03/23/2022 06:15:26",
      "content": "<p>Yes, and No, Fine tuning is providing some boost, but there are boosts because of TF-TPU is volatile in this competiton, as for me, its jumping here and there with little changes in the code. When using pytorch, its harder to get the score on par with TF-TPU, to learn how top LB is working, you dont need to wait for the competition to finish, you can just read other finished competitions discussions and find similarities. See the difference, what makes a top5 score, top10 score and top20 or whatever level you want to take a look at.</p>",
      "rawMarkdown": "Yes, and No, Fine tuning is providing some boost, but there are boosts because of TF-TPU is volatile in this competiton, as for me, its jumping here and there with little changes in the code. When using pytorch, its harder to get the score on par with TF-TPU, to learn how top LB is working, you dont need to wait for the competition to finish, you can just read other finished competitions discussions and find similarities. See the difference, what makes a top5 score, top10 score and top20 or whatever level you want to take a look at.",
      "votes": null
    },
    {
      "id": "1732255",
      "postDate": "03/23/2022 07:04:07",
      "content": "<p>Thanks for your suggestion</p>",
      "rawMarkdown": "Thanks for your suggestion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1731685,
      "author_name": "ptran1203",
      "author_url": "",
      "post_date": "03/22/2022 15:45:40",
      "content": "<p>You can read top solutions when the competition end.<br>\nTo improve LB there is a combination of many things: Dataset, Data augmentation, model architecture, ensembling, …</p>",
      "votes": null,
      "replies": [
        {
          "id": 1732058,
          "author_name": "bitlcc",
          "author_url": "",
          "post_date": "03/23/2022 00:50:42",
          "content": "<p>But In this competition, I see many kagglers use the same Dataset, Data augmentation and model architecture, but scores various from 0.2 to 0.7 plus. Ensembling can't  imporve so much, so other improvement can all attribute to fine-tuning?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732219,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/23/2022 06:15:26",
          "content": "<p>Yes, and No, Fine tuning is providing some boost, but there are boosts because of TF-TPU is volatile in this competiton, as for me, its jumping here and there with little changes in the code. When using pytorch, its harder to get the score on par with TF-TPU, to learn how top LB is working, you dont need to wait for the competition to finish, you can just read other finished competitions discussions and find similarities. See the difference, what makes a top5 score, top10 score and top20 or whatever level you want to take a look at.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732255,
          "author_name": "bitlcc",
          "author_url": "",
          "post_date": "03/23/2022 07:04:07",
          "content": "<p>Thanks for your suggestion</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1731642": "# Q1\nI wonder **how I can improve my LB scores in this competition**. No need to be too specific if it's not convenient. Just some direction is enough.\nMy confusion is most of the kagglers on top use the same model(as I see in public notebook), do they have special techniques **except fine-tuning the parameters**(or sharing some useful tricks on fine-tuning will be also grateful)?\n# Q2\nIn this competition, I see pretraining a good embedding and then use KNN is much better than directly classification.\n I have a test offline, even I only classfiy  a species with 1,000 unique individuals, the MAP@5 is always no more than 0.2, obviously far more than the SOTA . Maybe it's partially because I didn't do any fine-tuning, but I wonder **why the former method is much better than the latter**?\n\n**Previously thanks for your help.**",
    "1731685": "You can read top solutions when the competition end.\nTo improve LB there is a combination of many things: Dataset, Data augmentation, model architecture, ensembling, ...",
    "1732058": "But In this competition, I see many kagglers use the same Dataset, Data augmentation and model architecture, but scores various from 0.2 to 0.7 plus. Ensembling can't  imporve so much, so other improvement can all attribute to fine-tuning?",
    "1732219": "Yes, and No, Fine tuning is providing some boost, but there are boosts because of TF-TPU is volatile in this competiton, as for me, its jumping here and there with little changes in the code. When using pytorch, its harder to get the score on par with TF-TPU, to learn how top LB is working, you dont need to wait for the competition to finish, you can just read other finished competitions discussions and find similarities. See the difference, what makes a top5 score, top10 score and top20 or whatever level you want to take a look at.",
    "1732255": "Thanks for your suggestion"
  },
  "source": "meta"
}