{
  "id": 319789,
  "title": "3rd solution【Part】",
  "url": "/competitions/happy-whale-and-dolphin/discussion/319789",
  "author_name": "老肥",
  "post_date": "2022-04-19T00:56:31.097000",
  "votes": 61,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.</p>\n<p>Thanks for great teammates <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> 、@tikboa 、@qiuyanxinjupiter、@cuteffff .  Congratulations to us!</p>\n<p>Total solution can be found <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319896\" target=\"_blank\">here</a>, and this is our part solution. </p>\n<ol>\n<li>Brief. In our part solution, we use this wonderful notebook <a href=\"https://www.kaggle.com/code/jpbremer/backfin-detection-with-yolov5\" target=\"_blank\">backfin-detection-with-yolov5</a> to get data for part. The main idea is model can be easier to learn from pictures with backfin(seems as part), it can  be used for ensembling with body，and substitute the predictions of all whales without a backfin in the end.</li>\n<li>Model. We use effv1-b7 with Arcface and resolution of 512x512 for model 1, it has cv around 880, use effv1-b7 with CurricularFace + bnneck and resolution of 512x512 for model 2, it has cv around 880 too, training notebook can be found <a href=\"https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-curri-fold1\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-arc-fold1\" target=\"_blank\">here</a>, nothing magic.</li>\n<li>Train. We use kaggle TPU resource to train above model for 3folds of 10folds random split, pseudo label is used from ensemble model predicting the test set. </li>\n<li>Ensemble. For each fold, we concat 2 part model  and 6 body model embddings for total 8 * 512=4096 dimension to calculate distance and select the nearst results, we also use new_id replace for postprocessing. Notice that, before concat embedding, we need to normalize the embdding for each 512 dimension internally. We get around 0.930 local cv for ensembling. </li>\n<li>Optimize. We use bayesian optimization for finding the best weight of each 512 embedding，It can also works in private LB.</li>\n</ol>\n<p>Thanks for your reading!</p>",
  "messages": [
    {
      "id": 1759857,
      "postDate": "2022-04-19T00:56:31.097Z",
      "content": "<p>Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.</p>\n<p>Thanks for great teammates <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> 、@tikboa 、@qiuyanxinjupiter、@cuteffff .  Congratulations to us!</p>\n<p>Total solution can be found <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319896\" target=\"_blank\">here</a>, and this is our part solution. </p>\n<ol>\n<li>Brief. In our part solution, we use this wonderful notebook <a href=\"https://www.kaggle.com/code/jpbremer/backfin-detection-with-yolov5\" target=\"_blank\">backfin-detection-with-yolov5</a> to get data for part. The main idea is model can be easier to learn from pictures with backfin(seems as part), it can  be used for ensembling with body，and substitute the predictions of all whales without a backfin in the end.</li>\n<li>Model. We use effv1-b7 with Arcface and resolution of 512x512 for model 1, it has cv around 880, use effv1-b7 with CurricularFace + bnneck and resolution of 512x512 for model 2, it has cv around 880 too, training notebook can be found <a href=\"https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-curri-fold1\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-arc-fold1\" target=\"_blank\">here</a>, nothing magic.</li>\n<li>Train. We use kaggle TPU resource to train above model for 3folds of 10folds random split, pseudo label is used from ensemble model predicting the test set. </li>\n<li>Ensemble. For each fold, we concat 2 part model  and 6 body model embddings for total 8 * 512=4096 dimension to calculate distance and select the nearst results, we also use new_id replace for postprocessing. Notice that, before concat embedding, we need to normalize the embdding for each 512 dimension internally. We get around 0.930 local cv for ensembling. </li>\n<li>Optimize. We use bayesian optimization for finding the best weight of each 512 embedding，It can also works in private LB.</li>\n</ol>\n<p>Thanks for your reading!</p>",
      "rawMarkdown": "Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.\n\nThanks for great teammates @biglafe 、@tikboa 、@qiuyanxinjupiter、@cuteffff .  Congratulations to us!\n \nTotal solution can be found [here](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319896), and this is our part solution. \n\n1. Brief. In our part solution, we use this wonderful notebook [backfin-detection-with-yolov5](https://www.kaggle.com/code/jpbremer/backfin-detection-with-yolov5) to get data for part. The main idea is model can be easier to learn from pictures with backfin(seems as part), it can  be used for ensembling with body，and substitute the predictions of all whales without a backfin in the end.\n2. Model. We use effv1-b7 with Arcface and resolution of 512x512 for model 1, it has cv around 880, use effv1-b7 with CurricularFace + bnneck and resolution of 512x512 for model 2, it has cv around 880 too, training notebook can be found [here](https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-curri-fold1) and [here](https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-arc-fold1), nothing magic.\n3. Train. We use kaggle TPU resource to train above model for 3folds of 10folds random split, pseudo label is used from ensemble model predicting the test set. \n4. Ensemble. For each fold, we concat 2 part model  and 6 body model embddings for total 8 * 512=4096 dimension to calculate distance and select the nearst results, we also use new_id replace for postprocessing. Notice that, before concat embedding, we need to normalize the embdding for each 512 dimension internally. We get around 0.930 local cv for ensembling. \n5. Optimize. We use bayesian optimization for finding the best weight of each 512 embedding，It can also works in private LB.\n\nThanks for your reading!",
      "votes": 60
    },
    {
      "id": 1763701,
      "postDate": "2022-04-21T18:10:43.033Z",
      "content": "<p>It's interesting read about winner solution. Thanks for sharing.</p>",
      "rawMarkdown": "It's interesting read about winner solution. Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1759945,
      "postDate": "2022-04-19T02:14:12.527Z",
      "content": "<p>Congrats. I look forward to see your total solution :)</p>",
      "rawMarkdown": "Congrats. I look forward to see your total solution :)",
      "votes": 1,
      "replies": [
        {
          "id": 1759973,
          "postDate": "2022-04-19T02:31:55.187Z",
          "content": "<p>Thanks, you did a great job too！</p>",
          "rawMarkdown": "Thanks, you did a great job too！",
          "votes": 2
        }
      ]
    },
    {
      "id": 1759933,
      "postDate": "2022-04-19T02:05:42.360Z",
      "content": "<p>Good job, your backfin datasets seems to be really good, or you use the public one?</p>",
      "rawMarkdown": "Good job, your backfin datasets seems to be really good, or you use the public one?",
      "votes": 1,
      "replies": [
        {
          "id": 1759954,
          "postDate": "2022-04-19T02:22:44.293Z",
          "content": "<p>We just use the public one, of course we try to re-tag the backfin and train a new detetor but the cv score has not been improved..</p>",
          "rawMarkdown": "We just use the public one, of course we try to re-tag the backfin and train a new detetor but the cv score has not been improved..",
          "votes": 1
        },
        {
          "id": 1760126,
          "postDate": "2022-04-19T04:49:02.353Z",
          "content": "<p>I wonder what details in your training scheme achieve sush amazing cv score </p>",
          "rawMarkdown": "I wonder what details in your training scheme achieve sush amazing cv score ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1759902,
      "postDate": "2022-04-19T01:28:53.373Z",
      "content": "<p>There must some incredible magic in the training/architecture to reach CV 880! Can't wait to read the whole solution.</p>",
      "rawMarkdown": "There must some incredible magic in the training/architecture to reach CV 880! Can't wait to read the whole solution.",
      "votes": 2,
      "replies": [
        {
          "id": 1759971,
          "postDate": "2022-04-19T02:30:57.387Z",
          "content": "<p>Nothing magic, you will see in the pulic notebook soon.</p>",
          "rawMarkdown": "Nothing magic, you will see in the pulic notebook soon."
        }
      ]
    },
    {
      "id": 1759894,
      "postDate": "2022-04-19T01:20:58.307Z",
      "content": "<p>Thanks for your sharing. Can you explain a bit more about your training we used B7 but never reach cv 880. </p>",
      "rawMarkdown": "Thanks for your sharing. Can you explain a bit more about your training we used B7 but never reach cv 880. ",
      "votes": 2,
      "replies": [
        {
          "id": 1759934,
          "postDate": "2022-04-19T02:05:45.800Z",
          "content": "<p>Me too. <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a> ，can you give us more details about training? :)</p>",
          "rawMarkdown": "Me too. @librauee ，can you give us more details about training? :)"
        },
        {
          "id": 1759967,
          "postDate": "2022-04-19T02:29:50.560Z",
          "content": "<p>I will realease the training notebook soon😊, and my cv may different from yours(because our valid data is less than yours, just images with backfins).</p>",
          "rawMarkdown": "I will realease the training notebook soon😊, and my cv may different from yours(because our valid data is less than yours, just images with backfins).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1764516,
      "postDate": "2022-04-22T14:55:41.480Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a> amazing work! I just read your training notebook and it is really great source of learning for me. <br>\nIs any chance you make dataset public? Now it is trying to access private DS so it is hard to fully follow your way of thinking. Certainly I can replace with my one but … I will appreciate if you can share DS. </p>",
      "rawMarkdown": "Hi @librauee amazing work! I just read your training notebook and it is really great source of learning for me. \nIs any chance you make dataset public? Now it is trying to access private DS so it is hard to fully follow your way of thinking. Certainly I can replace with my one but ... I will appreciate if you can share DS. \n\n ",
      "replies": [
        {
          "id": 1764556,
          "postDate": "2022-04-22T15:26:08.887Z",
          "content": "<p>Of course， here are the dataset source：</p>\n<ol>\n<li>training and test dataset<br>\n<a href=\"https://www.kaggle.com/datasets/librauee/randomdatasetc\" target=\"_blank\">https://www.kaggle.com/datasets/librauee/randomdatasetc</a></li>\n<li>pseudo dataset<br>\n<a href=\"https://www.kaggle.com/datasets/librauee/randomdataseth\" target=\"_blank\">https://www.kaggle.com/datasets/librauee/randomdataseth</a></li>\n</ol>\n<p>and you can get the GCS path by the follows：</p>\n<p><code>GCS_PATH1 = KaggleDatasets().get_gcs_path('randomdatasetc')</code></p>",
          "rawMarkdown": "Of course， here are the dataset source：\n1. training and test dataset\nhttps://www.kaggle.com/datasets/librauee/randomdatasetc\n2. pseudo dataset\n https://www.kaggle.com/datasets/librauee/randomdataseth\n\nand you can get the GCS path by the follows：\n\n`GCS_PATH1 = KaggleDatasets().get_gcs_path('randomdatasetc')`"
        },
        {
          "id": 1764559,
          "postDate": "2022-04-22T15:28:05.457Z",
          "content": "<p>Many thanks <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a>. Reading your trick and playing with your notebook! Absolutely outstanding work. 👍👍👍</p>",
          "rawMarkdown": "Many thanks @librauee. Reading your trick and playing with your notebook! Absolutely outstanding work. 👍👍👍"
        }
      ]
    },
    {
      "id": 1761888,
      "postDate": "2022-04-20T09:21:12.727Z",
      "content": "<p>Congratulations to all …very insightful work…</p>",
      "rawMarkdown": "Congratulations to all ...very insightful work..."
    },
    {
      "id": 1761168,
      "postDate": "2022-04-19T17:57:19.113Z",
      "content": "<p>Congrats! I read your notebook and find you monitored Top5 Acc during training. May I confirm the model you get over 0.88lb has Top5 Acc around 0.64?</p>",
      "rawMarkdown": "Congrats! I read your notebook and find you monitored Top5 Acc during training. May I confirm the model you get over 0.88lb has Top5 Acc around 0.64?",
      "replies": [
        {
          "id": 1761420,
          "postDate": "2022-04-19T22:29:14.607Z",
          "content": "<p>Score 0.88 is not LB， just local cv(contains about 4k images with backfins).</p>",
          "rawMarkdown": "Score 0.88 is not LB， just local cv(contains about 4k images with backfins)."
        }
      ]
    },
    {
      "id": 1759943,
      "postDate": "2022-04-19T02:13:08.597Z",
      "content": "<p>can u explain about bayesian optimization or some source about it? and how u apply it for finding best weight?</p>",
      "rawMarkdown": "can u explain about bayesian optimization or some source about it? and how u apply it for finding best weight?",
      "replies": [
        {
          "id": 1759961,
          "postDate": "2022-04-19T02:27:38.910Z",
          "content": "<p>I use optuna， here is the <a href=\"https://optuna.org/\" target=\"_blank\">link</a>.</p>",
          "rawMarkdown": "I use optuna， here is the [link](https://optuna.org/)."
        }
      ]
    },
    {
      "id": 1759927,
      "postDate": "2022-04-19T02:01:18.613Z",
      "content": "<p>First of all, congratulations to you and <a href=\"https://www.kaggle.com/Cuteffff\" target=\"_blank\">@Cuteffff</a> for becoming kaggle masters.:D  I want to ask how many pseudo labels you use and how much it improves your model scores？</p>",
      "rawMarkdown": "First of all, congratulations to you and @Cuteffff for becoming kaggle masters.:D  I want to ask how many pseudo labels you use and how much it improves your model scores？",
      "replies": [
        {
          "id": 1759960,
          "postDate": "2022-04-19T02:25:41.843Z",
          "content": "<p>approximately 17000， and it can boost about 0.03-0.04.</p>",
          "rawMarkdown": "approximately 17000， and it can boost about 0.03-0.04."
        }
      ]
    },
    {
      "id": 1759922,
      "postDate": "2022-04-19T01:56:57.713Z",
      "content": "<p>恭喜！想问下，在分别对每个模型的embedding做normalize的时候，是在train的embedding上进行fit_transform，然后将得到的均值方差应用到test的 embedding上去吗？</p>",
      "rawMarkdown": "恭喜！想问下，在分别对每个模型的embedding做normalize的时候，是在train的embedding上进行fit_transform，然后将得到的均值方差应用到test的 embedding上去吗？",
      "replies": [
        {
          "id": 1759958,
          "postDate": "2022-04-19T02:24:15.793Z",
          "content": "<p>用的是训练集和测试集concat在一起做normalize </p>\n<pre><code>def normalize_list_numpy(list):\n    normalized_list = list / np.linalg.norm(list, ord=1)\n    return normalized_list\n</code></pre>",
          "rawMarkdown": "用的是训练集和测试集concat在一起做normalize \n\n```\ndef normalize_list_numpy(list):\n    normalized_list = list / np.linalg.norm(list, ord=1)\n    return normalized_list\n```",
          "votes": 2
        },
        {
          "id": 1759976,
          "postDate": "2022-04-19T02:32:14.667Z",
          "content": "<p>请问代码会放到kaggle或者您的github中吗，想学习一下😂</p>",
          "rawMarkdown": "请问代码会放到kaggle或者您的github中吗，想学习一下😂"
        }
      ]
    },
    {
      "id": 1759879,
      "postDate": "2022-04-19T01:12:26.173Z",
      "content": "<p>Now approach! Have you tried larger backbone like efficientnet v2 XL? As a starter, I wonder why people doesn't use it</p>",
      "rawMarkdown": "Now approach! Have you tried larger backbone like efficientnet v2 XL? As a starter, I wonder why people doesn't use it",
      "replies": [
        {
          "id": 1759889,
          "postDate": "2022-04-19T01:16:53.870Z",
          "content": "<p>We tried v2 XL with same dataset, the cv score is around 0.875.</p>",
          "rawMarkdown": "We tried v2 XL with same dataset, the cv score is around 0.875.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1759867,
      "postDate": "2022-04-19T01:02:12.263Z",
      "content": "<p>所以关键就是使用yolov5去crop出目标？😳</p>",
      "rawMarkdown": "所以关键就是使用yolov5去crop出目标？😳",
      "replies": [
        {
          "id": 1759888,
          "postDate": "2022-04-19T01:15:57.650Z",
          "content": "<p>对原始数据集的改进很重要🤓</p>",
          "rawMarkdown": "对原始数据集的改进很重要🤓",
          "votes": 1
        }
      ]
    },
    {
      "id": 2311523,
      "postDate": "2023-06-21T07:57:18.220Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2311781,
          "postDate": "2023-06-21T12:59:25.853Z",
          "content": "<p>用roberta、deberta啥的跑个0.96没问题呀，我主要是用lgb做的，做的特征可能好点👀</p>",
          "rawMarkdown": "用roberta、deberta啥的跑个0.96没问题呀，我主要是用lgb做的，做的特征可能好点👀",
          "replies": [
            {
              "id": 2312749,
              "postDate": "2023-06-22T07:02:04.750Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2315283,
              "postDate": "2023-06-24T02:57:50.940Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2324108,
              "postDate": "2023-06-30T11:43:54.447Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 1760623,
      "postDate": "2022-04-19T12:22:52.500Z",
      "content": "<p>Congrats, and thanks for sharing!</p>",
      "rawMarkdown": "Congrats, and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1759986,
      "postDate": "2022-04-19T02:41:46.903Z",
      "content": "<p>Congrats, and thanks for sharing!</p>",
      "rawMarkdown": "Congrats, and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1763337,
      "postDate": "2022-04-21T12:52:02.453Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!"
    },
    {
      "id": 1761247,
      "postDate": "2022-04-19T18:23:34.757Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 1763701,
      "author_name": "Alex Parkhomenko",
      "author_url": "",
      "post_date": "2022-04-21T18:10:43.033000",
      "content": "<p>It's interesting read about winner solution. Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1759945,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2022-04-19T02:14:12.527000",
      "content": "<p>Congrats. I look forward to see your total solution :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1759973,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:31:55.187000",
          "content": "<p>Thanks, you did a great job too！</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1759933,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2022-04-19T02:05:42.360000",
      "content": "<p>Good job, your backfin datasets seems to be really good, or you use the public one?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1759954,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:22:44.293000",
          "content": "<p>We just use the public one, of course we try to re-tag the backfin and train a new detetor but the cv score has not been improved..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1760126,
          "author_name": "Martin Kovacevic Buvinic",
          "author_url": "",
          "post_date": "2022-04-19T04:49:02.353000",
          "content": "<p>I wonder what details in your training scheme achieve sush amazing cv score </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1759902,
      "author_name": "Lex Toumbourou",
      "author_url": "",
      "post_date": "2022-04-19T01:28:53.373000",
      "content": "<p>There must some incredible magic in the training/architecture to reach CV 880! Can't wait to read the whole solution.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1759971,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:30:57.387000",
          "content": "<p>Nothing magic, you will see in the pulic notebook soon.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1759894,
      "author_name": "cuongnn",
      "author_url": "",
      "post_date": "2022-04-19T01:20:58.307000",
      "content": "<p>Thanks for your sharing. Can you explain a bit more about your training we used B7 but never reach cv 880. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1759934,
          "author_name": "Yangranran",
          "author_url": "",
          "post_date": "2022-04-19T02:05:45.800000",
          "content": "<p>Me too. <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a> ，can you give us more details about training? :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1759967,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:29:50.560000",
          "content": "<p>I will realease the training notebook soon😊, and my cv may different from yours(because our valid data is less than yours, just images with backfins).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1764516,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-04-22T14:55:41.480000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a> amazing work! I just read your training notebook and it is really great source of learning for me. <br>\nIs any chance you make dataset public? Now it is trying to access private DS so it is hard to fully follow your way of thinking. Certainly I can replace with my one but … I will appreciate if you can share DS. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1764556,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-22T15:26:08.887000",
          "content": "<p>Of course， here are the dataset source：</p>\n<ol>\n<li>training and test dataset<br>\n<a href=\"https://www.kaggle.com/datasets/librauee/randomdatasetc\" target=\"_blank\">https://www.kaggle.com/datasets/librauee/randomdatasetc</a></li>\n<li>pseudo dataset<br>\n<a href=\"https://www.kaggle.com/datasets/librauee/randomdataseth\" target=\"_blank\">https://www.kaggle.com/datasets/librauee/randomdataseth</a></li>\n</ol>\n<p>and you can get the GCS path by the follows：</p>\n<p><code>GCS_PATH1 = KaggleDatasets().get_gcs_path('randomdatasetc')</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1764559,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-22T15:28:05.457000",
          "content": "<p>Many thanks <a href=\"https://www.kaggle.com/librauee\" target=\"_blank\">@librauee</a>. Reading your trick and playing with your notebook! Absolutely outstanding work. 👍👍👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1761888,
      "author_name": "Iqbal Ali",
      "author_url": "",
      "post_date": "2022-04-20T09:21:12.727000",
      "content": "<p>Congratulations to all …very insightful work…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1761168,
      "author_name": "Zhongkai Shangguan",
      "author_url": "",
      "post_date": "2022-04-19T17:57:19.113000",
      "content": "<p>Congrats! I read your notebook and find you monitored Top5 Acc during training. May I confirm the model you get over 0.88lb has Top5 Acc around 0.64?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1761420,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T22:29:14.607000",
          "content": "<p>Score 0.88 is not LB， just local cv(contains about 4k images with backfins).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1759943,
      "author_name": "Mạnh Đỗ",
      "author_url": "",
      "post_date": "2022-04-19T02:13:08.597000",
      "content": "<p>can u explain about bayesian optimization or some source about it? and how u apply it for finding best weight?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1759961,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:27:38.910000",
          "content": "<p>I use optuna， here is the <a href=\"https://optuna.org/\" target=\"_blank\">link</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1759927,
      "author_name": "Yangranran",
      "author_url": "",
      "post_date": "2022-04-19T02:01:18.613000",
      "content": "<p>First of all, congratulations to you and <a href=\"https://www.kaggle.com/Cuteffff\" target=\"_blank\">@Cuteffff</a> for becoming kaggle masters.:D  I want to ask how many pseudo labels you use and how much it improves your model scores？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1759960,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:25:41.843000",
          "content": "<p>approximately 17000， and it can boost about 0.03-0.04.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1759922,
      "author_name": "Time Master",
      "author_url": "",
      "post_date": "2022-04-19T01:56:57.713000",
      "content": "<p>恭喜！想问下，在分别对每个模型的embedding做normalize的时候，是在train的embedding上进行fit_transform，然后将得到的均值方差应用到test的 embedding上去吗？</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1759958,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T02:24:15.793000",
          "content": "<p>用的是训练集和测试集concat在一起做normalize </p>\n<pre><code>def normalize_list_numpy(list):\n    normalized_list = list / np.linalg.norm(list, ord=1)\n    return normalized_list\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1759976,
          "author_name": "Yangranran",
          "author_url": "",
          "post_date": "2022-04-19T02:32:14.667000",
          "content": "<p>请问代码会放到kaggle或者您的github中吗，想学习一下😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1759879,
      "author_name": "Yeongjin Gwak",
      "author_url": "",
      "post_date": "2022-04-19T01:12:26.173000",
      "content": "<p>Now approach! Have you tried larger backbone like efficientnet v2 XL? As a starter, I wonder why people doesn't use it</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1759889,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T01:16:53.870000",
          "content": "<p>We tried v2 XL with same dataset, the cv score is around 0.875.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1759867,
      "author_name": "limzero",
      "author_url": "",
      "post_date": "2022-04-19T01:02:12.263000",
      "content": "<p>所以关键就是使用yolov5去crop出目标？😳</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1759888,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-04-19T01:15:57.650000",
          "content": "<p>对原始数据集的改进很重要🤓</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2311523,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-21T07:57:18.220000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2311781,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2023-06-21T12:59:25.853000",
          "content": "<p>用roberta、deberta啥的跑个0.96没问题呀，我主要是用lgb做的，做的特征可能好点👀</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2312749,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-22T07:02:04.750000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2315283,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-24T02:57:50.940000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2324108,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-30T11:43:54.447000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1760623,
      "author_name": "Antonio Martín",
      "author_url": "",
      "post_date": "2022-04-19T12:22:52.500000",
      "content": "<p>Congrats, and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1759986,
      "author_name": "aper",
      "author_url": "",
      "post_date": "2022-04-19T02:41:46.903000",
      "content": "<p>Congrats, and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763337,
      "author_name": "Flavio Cavalcante",
      "author_url": "",
      "post_date": "2022-04-21T12:52:02.453000",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1761247,
      "author_name": "Django",
      "author_url": "",
      "post_date": "2022-04-19T18:23:34.757000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1759857": "Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.\n\nThanks for great teammates @biglafe 、@tikboa 、@qiuyanxinjupiter、@cuteffff .  Congratulations to us!\n \nTotal solution can be found [here](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319896), and this is our part solution. \n\n1. Brief. In our part solution, we use this wonderful notebook [backfin-detection-with-yolov5](https://www.kaggle.com/code/jpbremer/backfin-detection-with-yolov5) to get data for part. The main idea is model can be easier to learn from pictures with backfin(seems as part), it can  be used for ensembling with body，and substitute the predictions of all whales without a backfin in the end.\n2. Model. We use effv1-b7 with Arcface and resolution of 512x512 for model 1, it has cv around 880, use effv1-b7 with CurricularFace + bnneck and resolution of 512x512 for model 2, it has cv around 880 too, training notebook can be found [here](https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-curri-fold1) and [here](https://www.kaggle.com/code/librauee/train-newpse-backfins-b7-512-arc-fold1), nothing magic.\n3. Train. We use kaggle TPU resource to train above model for 3folds of 10folds random split, pseudo label is used from ensemble model predicting the test set. \n4. Ensemble. For each fold, we concat 2 part model  and 6 body model embddings for total 8 * 512=4096 dimension to calculate distance and select the nearst results, we also use new_id replace for postprocessing. Notice that, before concat embedding, we need to normalize the embdding for each 512 dimension internally. We get around 0.930 local cv for ensembling. \n5. Optimize. We use bayesian optimization for finding the best weight of each 512 embedding，It can also works in private LB.\n\nThanks for your reading!",
    "1763701": "It's interesting read about winner solution. Thanks for sharing.",
    "1759945": "Congrats. I look forward to see your total solution :)",
    "1759933": "Good job, your backfin datasets seems to be really good, or you use the public one?",
    "1759902": "There must some incredible magic in the training/architecture to reach CV 880! Can't wait to read the whole solution.",
    "1759894": "Thanks for your sharing. Can you explain a bit more about your training we used B7 but never reach cv 880. ",
    "1764516": "Hi @librauee amazing work! I just read your training notebook and it is really great source of learning for me. \nIs any chance you make dataset public? Now it is trying to access private DS so it is hard to fully follow your way of thinking. Certainly I can replace with my one but ... I will appreciate if you can share DS. \n\n ",
    "1761888": "Congratulations to all ...very insightful work...",
    "1761168": "Congrats! I read your notebook and find you monitored Top5 Acc during training. May I confirm the model you get over 0.88lb has Top5 Acc around 0.64?",
    "1759943": "can u explain about bayesian optimization or some source about it? and how u apply it for finding best weight?",
    "1759927": "First of all, congratulations to you and @Cuteffff for becoming kaggle masters.:D  I want to ask how many pseudo labels you use and how much it improves your model scores？",
    "1759922": "恭喜！想问下，在分别对每个模型的embedding做normalize的时候，是在train的embedding上进行fit_transform，然后将得到的均值方差应用到test的 embedding上去吗？",
    "1759879": "Now approach! Have you tried larger backbone like efficientnet v2 XL? As a starter, I wonder why people doesn't use it",
    "1759867": "所以关键就是使用yolov5去crop出目标？😳",
    "2311523": "",
    "1760623": "Congrats, and thanks for sharing!",
    "1759986": "Congrats, and thanks for sharing!",
    "1763337": "Thanks for sharing!!",
    "1761247": "Thanks for sharing!"
  }
}