{
  "id": 308991,
  "title": "[SHORTCUT] Competition logbook - updated everyday",
  "url": "/competitions/happy-whale-and-dolphin/discussion/308991",
  "author_name": "",
  "post_date": "2022-02-21T10:40:52.256851100Z",
  "votes": 93,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Traditionally, where I take part in the competition, I try to browse the discussion forum and notebooks every day and make a short summary of the news. Kagglers who are just entering the competition will be able to quickly find the most important (in my subjective opinion) topics of the competition. If you have other suggestions, please report such threads. I will try to filter out the most important works that have an impact on the course of the competition.</p>\n<p>Let's start! </p>\n<p><strong>2022.02.21 (topic creation and important topics so far)</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.784  (2) 0.766 (3) 0.756</li>\n<li><strong>Best LB Score</strong> (<strong>0.679</strong>) solution so far is <a href=\"https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\" target=\"_blank\">Happywhale - Effnet B7 fork with Detic crop</a>by <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> . As we can see, the improvement of the result may be caused by the use of a network with more parameters B7 instead of B6 or DTIC crop (reduction of noise data).</li>\n<li><strong>Dataset</strong>- Description of problematic group (photos divided into groups) - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026\" target=\"_blank\">Things to know before starting image preprocessing</a>. Overview of dataset and problems that <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> found during data analyzis.</li>\n<li><strong>Dataset</strong>- Detic in <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503\" target=\"_blank\">cropped&amp;resized(512x512) dataset using detic</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> - in this topic you will find information how to deal with dataset (photos consists of many additional information so to reduce noise) - object detection (ROI) and cropping is proposal.</li>\n<li><strong>Dataset descriptive part</strong> - some fix to image description in file is needed. <a href=\"https://www.kaggle.com/karthickp6\" target=\"_blank\">@karthickp6</a> found and reported problem in <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304633\" target=\"_blank\">Duplicate names in species, can be merged together</a>. You need to fix some species names. </li>\n<li><strong>Competition metric explanations</strong> - in <a href=\"https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric\" target=\"_blank\">Explanation of MAP5 scoring metric</a> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> explained competition metric which is MAP5 (Mean Average Precision). </li>\n<li><strong>Cross Validation Strategy</strong>- how to validate our solution is described in <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> topic <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/306521\" target=\"_blank\">Stratified KFold v. Group KFold (aka. I'm a dummy)</a></li>\n<li><strong>Pipeline</strong> - one of the proposed solution pipeline was described in topic <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308573\" target=\"_blank\">Starting with domain knowledge and FIN-PRINT - inspiration how to approach problem</a> - FIN-PRINT</li>\n<li><strong>Tips</strong> - how to deal with large-scale recognition? <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> provided very interesting post about  Google Landmark Recognition 2020 Kaggle competition and key takeaways in area of large-scale recognition. <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304966\" target=\"_blank\">Grandmaster Series - How to Perform Large-Scale Image Classification</a></li>\n<li><strong>Starter notebooks</strong> - my favorite one is <a href=\"https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter\" target=\"_blank\">[Pytorch] ArcFace + GeM Pooling Starter</a> by <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a>. It contains great set of techniques can be useful in this competition (implementation in Pytorch)-&gt; ArcFace, GeM Pooling, nice shedule switch and w&amp;b training logging integration. Second one by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> is <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">HappyWhale ArcFace Baseline (TPU)</a>. It is implemented in TensorFlow (TPU) + ArcFace, EfficientNet as a backbone/feature extractor.</li>\n<li><strong>Previous competition</strong> - solution from previous competition collected <a href=\"https://www.kaggle.com/vad13irt\" target=\"_blank\">@vad13irt</a> in topic <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304504\" target=\"_blank\">Previous Happywhale Competition Solutions</a></li>\n</ul>\n<p><strong>2022.02.22</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.810</strong> (2) 0.770 (3) 0.766 - first score above 0.8 (!) by <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> - congratulations!</li>\n<li><strong>Tutorial</strong> - really interesting tutorial by <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> covering many important topics (all in one) - EDA, TPU usage, TF-Records, Data Augumentations, <strong>Involutional Neural Networks</strong> I have not heard about this type of NN so far - interesting! -&gt; <a href=\"https://www.kaggle.com/usharengaraju/tensorflow-tpu-involutionalneuralnetworks/comments\" target=\"_blank\">[TensorFlow,TPU]InvolutionalNeuralNetworks</a></li>\n<li><strong>Dataset</strong> - new dataset attitude - generate crops of dorsal fins and remove background - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214\" target=\"_blank\">DATASET - dorsal fins for all IDs without background</a> </li>\n</ul>\n<p><strong>2022.02.23</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.770 (3) 0.766 - no changes </li>\n<li><strong>Notebook - JAX/FLAX</strong> - Jax is in my area of interest - I believe it will soon be one of the most significant programming languages ​​in the ML area. It's worth studying. <a href=\"https://www.kaggle.com/alexlwh\" target=\"_blank\">@alexlwh</a> show us not only JAX but NN library FLAX. Notebook is really great! <a href=\"https://www.kaggle.com/alexlwh/happywhale-flax-jax-resnet-baseline\" target=\"_blank\">HappyWhale 🔥Flax/JAX - ResNet Baseline</a>   </li>\n<li><strong>Notebook</strong> - I really like notebook provided by <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> <a href=\"https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-train-rapids-clusters\" target=\"_blank\">Whales&amp;Dolphins: EffNet Train &amp; RAPIDS Clusters</a>. Why? She summarized the most important solutions in the area of large-scale class classification. In notebook you will find: Generalized Mean (GeM), Additive Angular Margin Loss, RAPIDS clustering. Absolutely great notebook written in Pytorch.</li>\n</ul>\n<p><strong>2022.02.24</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.771 (3) 0.766 </li>\n</ul>\n<p><strong>2022.02.25</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.771 (3) <strong>0.771</strong></li>\n<li><strong>Dataset</strong> - summary of all dataset created in this competition so far - topic created by <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309691\" target=\"_blank\">Dataset Dataset Dataset</a></li>\n</ul>\n<p><strong>2022.02.28</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.796 (3) <strong>0.792</strong></li>\n<li><strong>Best LB Score</strong> (<strong>0.729</strong>) solution so far is <a href=\"https://www.kaggle.com/aikhmelnytskyy/happywhale-arcface-baseline-eff7-tpu-768-concat\" target=\"_blank\">happywhale_arcface_baseline_eff7_tpu_768_concat</a>by <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a>. </li>\n<li><strong>Dataset</strong> - dorsal finds database created using yolov5 - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/310153\" target=\"_blank\">Releasing my Dorsal Fin Dataset &amp; Code</a></li>\n<li><strong>Notebook - Keypoint - LoFTR</strong> - <a href=\"https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia\" target=\"_blank\">Whales feature matching LoFTR - Kornia</a></li>\n</ul>\n<p><strong>2022.03.02</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.830 (2) 0.813 (3) 0.808</strong></li>\n<li>I can't see any new factors which can improve score so far</li>\n</ul>\n<p><strong>2022.03.08</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.840 (2) 0.835 (3) 0.830</strong></li>\n<li><strong>Notebook</strong> - <a href=\"https://www.kaggle.com/remekkinas/remove-background-salient-object-detection\" target=\"_blank\">Remove background - Salient Object Detection</a></li>\n</ul>\n<p><strong>2022.03.11</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.840 (2) 0.838 (3) 0.830</strong></li>\n<li><strong>Discussion</strong> - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/311887#1718683\" target=\"_blank\">How much does score gets better with bigger models and bigger image sizes?</a> by <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> - size score relationship </li>\n</ul>\n<p><strong>2022.03.14</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.844 (2) 0.840 (3) 0.834</strong></li>\n<li><strong>Discussion</strong> - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/312499\" target=\"_blank\">DATASET - Background removed 512x512 Happywhale dataset using State of the Art Salient Object Detector</a> by <a href=\"https://www.kaggle.com/adnanpen\" target=\"_blank\">@adnanpen</a> - another way to remove background </li>\n</ul>",
  "messages": [
    {
      "id": "1699655",
      "postDate": "02/21/2022 10:40:52",
      "content": "<p>Traditionally, where I take part in the competition, I try to browse the discussion forum and notebooks every day and make a short summary of the news. Kagglers who are just entering the competition will be able to quickly find the most important (in my subjective opinion) topics of the competition. If you have other suggestions, please report such threads. I will try to filter out the most important works that have an impact on the course of the competition.</p>\n<p>Let's start! </p>\n<p><strong>2022.02.21 (topic creation and important topics so far)</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.784  (2) 0.766 (3) 0.756</li>\n<li><strong>Best LB Score</strong> (<strong>0.679</strong>) solution so far is <a href=\"https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop\" target=\"_blank\">Happywhale - Effnet B7 fork with Detic crop</a>by <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> . As we can see, the improvement of the result may be caused by the use of a network with more parameters B7 instead of B6 or DTIC crop (reduction of noise data).</li>\n<li><strong>Dataset</strong>- Description of problematic group (photos divided into groups) - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026\" target=\"_blank\">Things to know before starting image preprocessing</a>. Overview of dataset and problems that <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> found during data analyzis.</li>\n<li><strong>Dataset</strong>- Detic in <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503\" target=\"_blank\">cropped&amp;resized(512x512) dataset using detic</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> - in this topic you will find information how to deal with dataset (photos consists of many additional information so to reduce noise) - object detection (ROI) and cropping is proposal.</li>\n<li><strong>Dataset descriptive part</strong> - some fix to image description in file is needed. <a href=\"https://www.kaggle.com/karthickp6\" target=\"_blank\">@karthickp6</a> found and reported problem in <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304633\" target=\"_blank\">Duplicate names in species, can be merged together</a>. You need to fix some species names. </li>\n<li><strong>Competition metric explanations</strong> - in <a href=\"https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric\" target=\"_blank\">Explanation of MAP5 scoring metric</a> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> explained competition metric which is MAP5 (Mean Average Precision). </li>\n<li><strong>Cross Validation Strategy</strong>- how to validate our solution is described in <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> topic <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/306521\" target=\"_blank\">Stratified KFold v. Group KFold (aka. I'm a dummy)</a></li>\n<li><strong>Pipeline</strong> - one of the proposed solution pipeline was described in topic <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308573\" target=\"_blank\">Starting with domain knowledge and FIN-PRINT - inspiration how to approach problem</a> - FIN-PRINT</li>\n<li><strong>Tips</strong> - how to deal with large-scale recognition? <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> provided very interesting post about  Google Landmark Recognition 2020 Kaggle competition and key takeaways in area of large-scale recognition. <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304966\" target=\"_blank\">Grandmaster Series - How to Perform Large-Scale Image Classification</a></li>\n<li><strong>Starter notebooks</strong> - my favorite one is <a href=\"https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter\" target=\"_blank\">[Pytorch] ArcFace + GeM Pooling Starter</a> by <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a>. It contains great set of techniques can be useful in this competition (implementation in Pytorch)-&gt; ArcFace, GeM Pooling, nice shedule switch and w&amp;b training logging integration. Second one by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> is <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">HappyWhale ArcFace Baseline (TPU)</a>. It is implemented in TensorFlow (TPU) + ArcFace, EfficientNet as a backbone/feature extractor.</li>\n<li><strong>Previous competition</strong> - solution from previous competition collected <a href=\"https://www.kaggle.com/vad13irt\" target=\"_blank\">@vad13irt</a> in topic <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304504\" target=\"_blank\">Previous Happywhale Competition Solutions</a></li>\n</ul>\n<p><strong>2022.02.22</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.810</strong> (2) 0.770 (3) 0.766 - first score above 0.8 (!) by <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> - congratulations!</li>\n<li><strong>Tutorial</strong> - really interesting tutorial by <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> covering many important topics (all in one) - EDA, TPU usage, TF-Records, Data Augumentations, <strong>Involutional Neural Networks</strong> I have not heard about this type of NN so far - interesting! -&gt; <a href=\"https://www.kaggle.com/usharengaraju/tensorflow-tpu-involutionalneuralnetworks/comments\" target=\"_blank\">[TensorFlow,TPU]InvolutionalNeuralNetworks</a></li>\n<li><strong>Dataset</strong> - new dataset attitude - generate crops of dorsal fins and remove background - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214\" target=\"_blank\">DATASET - dorsal fins for all IDs without background</a> </li>\n</ul>\n<p><strong>2022.02.23</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.770 (3) 0.766 - no changes </li>\n<li><strong>Notebook - JAX/FLAX</strong> - Jax is in my area of interest - I believe it will soon be one of the most significant programming languages ​​in the ML area. It's worth studying. <a href=\"https://www.kaggle.com/alexlwh\" target=\"_blank\">@alexlwh</a> show us not only JAX but NN library FLAX. Notebook is really great! <a href=\"https://www.kaggle.com/alexlwh/happywhale-flax-jax-resnet-baseline\" target=\"_blank\">HappyWhale 🔥Flax/JAX - ResNet Baseline</a>   </li>\n<li><strong>Notebook</strong> - I really like notebook provided by <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> <a href=\"https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-train-rapids-clusters\" target=\"_blank\">Whales&amp;Dolphins: EffNet Train &amp; RAPIDS Clusters</a>. Why? She summarized the most important solutions in the area of large-scale class classification. In notebook you will find: Generalized Mean (GeM), Additive Angular Margin Loss, RAPIDS clustering. Absolutely great notebook written in Pytorch.</li>\n</ul>\n<p><strong>2022.02.24</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.771 (3) 0.766 </li>\n</ul>\n<p><strong>2022.02.25</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.771 (3) <strong>0.771</strong></li>\n<li><strong>Dataset</strong> - summary of all dataset created in this competition so far - topic created by <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309691\" target=\"_blank\">Dataset Dataset Dataset</a></li>\n</ul>\n<p><strong>2022.02.28</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - (1) 0.810 (2) 0.796 (3) <strong>0.792</strong></li>\n<li><strong>Best LB Score</strong> (<strong>0.729</strong>) solution so far is <a href=\"https://www.kaggle.com/aikhmelnytskyy/happywhale-arcface-baseline-eff7-tpu-768-concat\" target=\"_blank\">happywhale_arcface_baseline_eff7_tpu_768_concat</a>by <a href=\"https://www.kaggle.com/aikhmelnytskyy\" target=\"_blank\">@aikhmelnytskyy</a>. </li>\n<li><strong>Dataset</strong> - dorsal finds database created using yolov5 - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/310153\" target=\"_blank\">Releasing my Dorsal Fin Dataset &amp; Code</a></li>\n<li><strong>Notebook - Keypoint - LoFTR</strong> - <a href=\"https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia\" target=\"_blank\">Whales feature matching LoFTR - Kornia</a></li>\n</ul>\n<p><strong>2022.03.02</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.830 (2) 0.813 (3) 0.808</strong></li>\n<li>I can't see any new factors which can improve score so far</li>\n</ul>\n<p><strong>2022.03.08</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.840 (2) 0.835 (3) 0.830</strong></li>\n<li><strong>Notebook</strong> - <a href=\"https://www.kaggle.com/remekkinas/remove-background-salient-object-detection\" target=\"_blank\">Remove background - Salient Object Detection</a></li>\n</ul>\n<p><strong>2022.03.11</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.840 (2) 0.838 (3) 0.830</strong></li>\n<li><strong>Discussion</strong> - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/311887#1718683\" target=\"_blank\">How much does score gets better with bigger models and bigger image sizes?</a> by <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> - size score relationship </li>\n</ul>\n<p><strong>2022.03.14</strong></p>\n<ul>\n<li><strong>The best LB score (TOP3)</strong> - <strong>(1) 0.844 (2) 0.840 (3) 0.834</strong></li>\n<li><strong>Discussion</strong> - <a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/312499\" target=\"_blank\">DATASET - Background removed 512x512 Happywhale dataset using State of the Art Salient Object Detector</a> by <a href=\"https://www.kaggle.com/adnanpen\" target=\"_blank\">@adnanpen</a> - another way to remove background </li>\n</ul>",
      "rawMarkdown": "Traditionally, where I take part in the competition, I try to browse the discussion forum and notebooks every day and make a short summary of the news. Kagglers who are just entering the competition will be able to quickly find the most important (in my subjective opinion) topics of the competition. If you have other suggestions, please report such threads. I will try to filter out the most important works that have an impact on the course of the competition.\n\nLet's start! \n\n**2022.02.21 (topic creation and important topics so far)**\n\n- **The best LB score (TOP3)** - (1) 0.784  (2) 0.766 (3) 0.756\n- **Best LB Score** (**0.679**) solution so far is [Happywhale - Effnet B7 fork with Detic crop](https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop)by @dragonzhang . As we can see, the improvement of the result may be caused by the use of a network with more parameters B7 instead of B6 or DTIC crop (reduction of noise data).\n- **Dataset**- Description of problematic group (photos divided into groups) - [Things to know before starting image preprocessing](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026). Overview of dataset and problems that @andradaolteanu found during data analyzis.\n- **Dataset**- Detic in [cropped&resized(512x512) dataset using detic](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503) by @phalanx - in this topic you will find information how to deal with dataset (photos consists of many additional information so to reduce noise) - object detection (ROI) and cropping is proposal.\n- **Dataset descriptive part** - some fix to image description in file is needed. @karthickp6 found and reported problem in [Duplicate names in species, can be merged together](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304633). You need to fix some species names. \n- **Competition metric explanations** - in [Explanation of MAP5 scoring metric](https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric) @pestipeti explained competition metric which is MAP5 (Mean Average Precision). \n- **Cross Validation Strategy**- how to validate our solution is described in @dschettler8845 topic [Stratified KFold v. Group KFold (aka. I'm a dummy)](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/306521)\n- **Pipeline** - one of the proposed solution pipeline was described in topic [Starting with domain knowledge and FIN-PRINT - inspiration how to approach problem](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308573) - FIN-PRINT\n- **Tips** - how to deal with large-scale recognition? @debarshichanda provided very interesting post about  Google Landmark Recognition 2020 Kaggle competition and key takeaways in area of large-scale recognition. [Grandmaster Series - How to Perform Large-Scale Image Classification](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304966)\n- **Starter notebooks** - my favorite one is [[Pytorch] ArcFace + GeM Pooling Starter](https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter) by @debarshichanda. It contains great set of techniques can be useful in this competition (implementation in Pytorch)-> ArcFace, GeM Pooling, nice shedule switch and w&b training logging integration. Second one by @ks2019 is [HappyWhale ArcFace Baseline (TPU)](https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu). It is implemented in TensorFlow (TPU) + ArcFace, EfficientNet as a backbone/feature extractor.\n- **Previous competition** - solution from previous competition collected @vad13irt in topic [Previous Happywhale Competition Solutions](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304504)\n\n**2022.02.22**\n- **The best LB score (TOP3)** - **(1) 0.810** (2) 0.770 (3) 0.766 - first score above 0.8 (!) by @biglafe - congratulations!\n- **Tutorial** - really interesting tutorial by @usharengaraju covering many important topics (all in one) - EDA, TPU usage, TF-Records, Data Augumentations, **Involutional Neural Networks** I have not heard about this type of NN so far - interesting! -> [[TensorFlow,TPU]InvolutionalNeuralNetworks](https://www.kaggle.com/usharengaraju/tensorflow-tpu-involutionalneuralnetworks/comments)\n- **Dataset** - new dataset attitude - generate crops of dorsal fins and remove background - [DATASET - dorsal fins for all IDs without background](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214) \n\n**2022.02.23**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.770 (3) 0.766 - no changes \n- **Notebook - JAX/FLAX** - Jax is in my area of interest - I believe it will soon be one of the most significant programming languages ​​in the ML area. It's worth studying. @alexlwh show us not only JAX but NN library FLAX. Notebook is really great! [HappyWhale 🔥Flax/JAX - ResNet Baseline](https://www.kaggle.com/alexlwh/happywhale-flax-jax-resnet-baseline)   \n- **Notebook** - I really like notebook provided by @andradaolteanu [Whales&Dolphins: EffNet Train & RAPIDS Clusters](https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-train-rapids-clusters). Why? She summarized the most important solutions in the area of large-scale class classification. In notebook you will find: Generalized Mean (GeM), Additive Angular Margin Loss, RAPIDS clustering. Absolutely great notebook written in Pytorch.\n\n**2022.02.24**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.771 (3) 0.766 \n\n**2022.02.25**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.771 (3) **0.771**\n- **Dataset** - summary of all dataset created in this competition so far - topic created by @ayuraj - [Dataset Dataset Dataset](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309691)\n\n**2022.02.28**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.796 (3) **0.792**\n- **Best LB Score** (**0.729**) solution so far is [happywhale_arcface_baseline_eff7_tpu_768_concat](https://www.kaggle.com/aikhmelnytskyy/happywhale-arcface-baseline-eff7-tpu-768-concat)by @aikhmelnytskyy. \n- **Dataset** - dorsal finds database created using yolov5 - [Releasing my Dorsal Fin Dataset & Code](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/310153)\n- **Notebook - Keypoint - LoFTR** - [Whales feature matching LoFTR - Kornia](https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia)\n\n**2022.03.02**\n- **The best LB score (TOP3)** - **(1) 0.830 (2) 0.813 (3) 0.808**\n- I can't see any new factors which can improve score so far\n\n**2022.03.08**\n- **The best LB score (TOP3)** - **(1) 0.840 (2) 0.835 (3) 0.830**\n- **Notebook** - [Remove background - Salient Object Detection](https://www.kaggle.com/remekkinas/remove-background-salient-object-detection)\n\n**2022.03.11**\n- **The best LB score (TOP3)** - **(1) 0.840 (2) 0.838 (3) 0.830**\n- **Discussion** - [How much does score gets better with bigger models and bigger image sizes?](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/311887#1718683) by @harshitsheoran - size score relationship \n\n**2022.03.14**\n- **The best LB score (TOP3)** - **(1) 0.844 (2) 0.840 (3) 0.834**\n- **Discussion** - [DATASET - Background removed 512x512 Happywhale dataset using State of the Art Salient Object Detector](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/312499) by @adnanpen - another way to remove background",
      "votes": null
    },
    {
      "id": "1699703",
      "postDate": "02/21/2022 11:22:40",
      "content": "<p>Thanks, maybe also do gold range scores for each entry, i.e 2022.02.21 .784 - .703 for some dynamics?</p>",
      "rawMarkdown": "Thanks, maybe also do gold range scores for each entry, i.e 2022.02.21 .784 - .703 for some dynamics?",
      "votes": null
    },
    {
      "id": "1699723",
      "postDate": "02/21/2022 11:47:35",
      "content": "<p>Good idea! :)</p>",
      "rawMarkdown": "Good idea! :)",
      "votes": null
    },
    {
      "id": "1699827",
      "postDate": "02/21/2022 13:22:02",
      "content": "<p>nice summary.  based on my a few experiments and comparison,  training batch size seems play important role for performance?</p>",
      "rawMarkdown": "nice summary.  based on my a few experiments and comparison,  training batch size seems play important role for performance?",
      "votes": null
    },
    {
      "id": "1700139",
      "postDate": "02/21/2022 17:31:32",
      "content": "<p>Thank you for information. I will post only experiments which are proven by experiment published in competition 🤓</p>",
      "rawMarkdown": "Thank you for information. I will post only experiments which are proven by experiment published in competition 🤓",
      "votes": null
    },
    {
      "id": "1701994",
      "postDate": "02/23/2022 08:26:50",
      "content": "<p>Updated. I found two absolutely fantastic notebook today.</p>",
      "rawMarkdown": "Updated. I found two absolutely fantastic notebook today.",
      "votes": null
    },
    {
      "id": "1707126",
      "postDate": "02/28/2022 06:39:23",
      "content": "<p>Updated 28.02.2022.</p>",
      "rawMarkdown": "Updated 28.02.2022.",
      "votes": null
    },
    {
      "id": "1718693",
      "postDate": "03/11/2022 04:42:11",
      "content": "<p>Dude take my updoot. I had all these notebooks scattered in bookmarks, but I really do appreciate you putting them all together into one place.</p>",
      "rawMarkdown": "Dude take my updoot. I had all these notebooks scattered in bookmarks, but I really do appreciate you putting them all together into one place.",
      "votes": null
    },
    {
      "id": "1718781",
      "postDate": "03/11/2022 06:42:59",
      "content": "<p>Thank you! At this time … only one section I update LB - score :) No new ideas which influence on score (in my opinion).</p>",
      "rawMarkdown": "Thank you! At this time ... only one section I update LB - score :) No new ideas which influence on score (in my opinion).",
      "votes": null
    },
    {
      "id": "1722007",
      "postDate": "03/14/2022 07:12:03",
      "content": "<p>I literally started this competition with this post. <br>\nThanks for your hard work, to help people.</p>",
      "rawMarkdown": "I literally started this competition with this post. \nThanks for your hard work, to help people.",
      "votes": null
    },
    {
      "id": "1722103",
      "postDate": "03/14/2022 08:28:56",
      "content": "<p>Great I am happy. Thank you! 👍🙏</p>",
      "rawMarkdown": "Great I am happy. Thank you! 👍🙏",
      "votes": null
    },
    {
      "id": "1732292",
      "postDate": "03/23/2022 07:42:02",
      "content": "<p>Thanks for ur nice organizing :)</p>",
      "rawMarkdown": "Thanks for ur nice organizing :)",
      "votes": null
    },
    {
      "id": "1732330",
      "postDate": "03/23/2022 08:56:13",
      "content": "<p>Update in progress :)</p>",
      "rawMarkdown": "Update in progress :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1699703,
      "author_name": "bakeryproducts",
      "author_url": "",
      "post_date": "02/21/2022 11:22:40",
      "content": "<p>Thanks, maybe also do gold range scores for each entry, i.e 2022.02.21 .784 - .703 for some dynamics?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1699723,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/21/2022 11:47:35",
          "content": "<p>Good idea! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1699827,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/21/2022 13:22:02",
      "content": "<p>nice summary.  based on my a few experiments and comparison,  training batch size seems play important role for performance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1700139,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/21/2022 17:31:32",
          "content": "<p>Thank you for information. I will post only experiments which are proven by experiment published in competition 🤓</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1701994,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/23/2022 08:26:50",
      "content": "<p>Updated. I found two absolutely fantastic notebook today.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1707126,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/28/2022 06:39:23",
      "content": "<p>Updated 28.02.2022.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1718693,
      "author_name": "dentistdad",
      "author_url": "",
      "post_date": "03/11/2022 04:42:11",
      "content": "<p>Dude take my updoot. I had all these notebooks scattered in bookmarks, but I really do appreciate you putting them all together into one place.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1718781,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/11/2022 06:42:59",
          "content": "<p>Thank you! At this time … only one section I update LB - score :) No new ideas which influence on score (in my opinion).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1722007,
      "author_name": "tharun2001",
      "author_url": "",
      "post_date": "03/14/2022 07:12:03",
      "content": "<p>I literally started this competition with this post. <br>\nThanks for your hard work, to help people.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1722103,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/14/2022 08:28:56",
          "content": "<p>Great I am happy. Thank you! 👍🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1732292,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "03/23/2022 07:42:02",
      "content": "<p>Thanks for ur nice organizing :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1732330,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/23/2022 08:56:13",
          "content": "<p>Update in progress :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1699655": "Traditionally, where I take part in the competition, I try to browse the discussion forum and notebooks every day and make a short summary of the news. Kagglers who are just entering the competition will be able to quickly find the most important (in my subjective opinion) topics of the competition. If you have other suggestions, please report such threads. I will try to filter out the most important works that have an impact on the course of the competition.\n\nLet's start! \n\n**2022.02.21 (topic creation and important topics so far)**\n\n- **The best LB score (TOP3)** - (1) 0.784  (2) 0.766 (3) 0.756\n- **Best LB Score** (**0.679**) solution so far is [Happywhale - Effnet B7 fork with Detic crop](https://www.kaggle.com/dragonzhang/happywhale-effnet-b7-fork-with-detic-crop)by @dragonzhang . As we can see, the improvement of the result may be caused by the use of a network with more parameters B7 instead of B6 or DTIC crop (reduction of noise data).\n- **Dataset**- Description of problematic group (photos divided into groups) - [Things to know before starting image preprocessing](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026). Overview of dataset and problems that @andradaolteanu found during data analyzis.\n- **Dataset**- Detic in [cropped&resized(512x512) dataset using detic](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503) by @phalanx - in this topic you will find information how to deal with dataset (photos consists of many additional information so to reduce noise) - object detection (ROI) and cropping is proposal.\n- **Dataset descriptive part** - some fix to image description in file is needed. @karthickp6 found and reported problem in [Duplicate names in species, can be merged together](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304633). You need to fix some species names. \n- **Competition metric explanations** - in [Explanation of MAP5 scoring metric](https://www.kaggle.com/pestipeti/explanation-of-map5-scoring-metric) @pestipeti explained competition metric which is MAP5 (Mean Average Precision). \n- **Cross Validation Strategy**- how to validate our solution is described in @dschettler8845 topic [Stratified KFold v. Group KFold (aka. I'm a dummy)](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/306521)\n- **Pipeline** - one of the proposed solution pipeline was described in topic [Starting with domain knowledge and FIN-PRINT - inspiration how to approach problem](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308573) - FIN-PRINT\n- **Tips** - how to deal with large-scale recognition? @debarshichanda provided very interesting post about  Google Landmark Recognition 2020 Kaggle competition and key takeaways in area of large-scale recognition. [Grandmaster Series - How to Perform Large-Scale Image Classification](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304966)\n- **Starter notebooks** - my favorite one is [[Pytorch] ArcFace + GeM Pooling Starter](https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter) by @debarshichanda. It contains great set of techniques can be useful in this competition (implementation in Pytorch)-> ArcFace, GeM Pooling, nice shedule switch and w&b training logging integration. Second one by @ks2019 is [HappyWhale ArcFace Baseline (TPU)](https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu). It is implemented in TensorFlow (TPU) + ArcFace, EfficientNet as a backbone/feature extractor.\n- **Previous competition** - solution from previous competition collected @vad13irt in topic [Previous Happywhale Competition Solutions](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304504)\n\n**2022.02.22**\n- **The best LB score (TOP3)** - **(1) 0.810** (2) 0.770 (3) 0.766 - first score above 0.8 (!) by @biglafe - congratulations!\n- **Tutorial** - really interesting tutorial by @usharengaraju covering many important topics (all in one) - EDA, TPU usage, TF-Records, Data Augumentations, **Involutional Neural Networks** I have not heard about this type of NN so far - interesting! -> [[TensorFlow,TPU]InvolutionalNeuralNetworks](https://www.kaggle.com/usharengaraju/tensorflow-tpu-involutionalneuralnetworks/comments)\n- **Dataset** - new dataset attitude - generate crops of dorsal fins and remove background - [DATASET - dorsal fins for all IDs without background](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214) \n\n**2022.02.23**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.770 (3) 0.766 - no changes \n- **Notebook - JAX/FLAX** - Jax is in my area of interest - I believe it will soon be one of the most significant programming languages ​​in the ML area. It's worth studying. @alexlwh show us not only JAX but NN library FLAX. Notebook is really great! [HappyWhale 🔥Flax/JAX - ResNet Baseline](https://www.kaggle.com/alexlwh/happywhale-flax-jax-resnet-baseline)   \n- **Notebook** - I really like notebook provided by @andradaolteanu [Whales&Dolphins: EffNet Train & RAPIDS Clusters](https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-train-rapids-clusters). Why? She summarized the most important solutions in the area of large-scale class classification. In notebook you will find: Generalized Mean (GeM), Additive Angular Margin Loss, RAPIDS clustering. Absolutely great notebook written in Pytorch.\n\n**2022.02.24**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.771 (3) 0.766 \n\n**2022.02.25**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.771 (3) **0.771**\n- **Dataset** - summary of all dataset created in this competition so far - topic created by @ayuraj - [Dataset Dataset Dataset](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309691)\n\n**2022.02.28**\n- **The best LB score (TOP3)** - (1) 0.810 (2) 0.796 (3) **0.792**\n- **Best LB Score** (**0.729**) solution so far is [happywhale_arcface_baseline_eff7_tpu_768_concat](https://www.kaggle.com/aikhmelnytskyy/happywhale-arcface-baseline-eff7-tpu-768-concat)by @aikhmelnytskyy. \n- **Dataset** - dorsal finds database created using yolov5 - [Releasing my Dorsal Fin Dataset & Code](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/310153)\n- **Notebook - Keypoint - LoFTR** - [Whales feature matching LoFTR - Kornia](https://www.kaggle.com/remekkinas/whales-feature-matching-loftr-kornia)\n\n**2022.03.02**\n- **The best LB score (TOP3)** - **(1) 0.830 (2) 0.813 (3) 0.808**\n- I can't see any new factors which can improve score so far\n\n**2022.03.08**\n- **The best LB score (TOP3)** - **(1) 0.840 (2) 0.835 (3) 0.830**\n- **Notebook** - [Remove background - Salient Object Detection](https://www.kaggle.com/remekkinas/remove-background-salient-object-detection)\n\n**2022.03.11**\n- **The best LB score (TOP3)** - **(1) 0.840 (2) 0.838 (3) 0.830**\n- **Discussion** - [How much does score gets better with bigger models and bigger image sizes?](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/311887#1718683) by @harshitsheoran - size score relationship \n\n**2022.03.14**\n- **The best LB score (TOP3)** - **(1) 0.844 (2) 0.840 (3) 0.834**\n- **Discussion** - [DATASET - Background removed 512x512 Happywhale dataset using State of the Art Salient Object Detector](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/312499) by @adnanpen - another way to remove background",
    "1699703": "Thanks, maybe also do gold range scores for each entry, i.e 2022.02.21 .784 - .703 for some dynamics?",
    "1699723": "Good idea! :)",
    "1699827": "nice summary.  based on my a few experiments and comparison,  training batch size seems play important role for performance?",
    "1700139": "Thank you for information. I will post only experiments which are proven by experiment published in competition 🤓",
    "1701994": "Updated. I found two absolutely fantastic notebook today.",
    "1707126": "Updated 28.02.2022.",
    "1718693": "Dude take my updoot. I had all these notebooks scattered in bookmarks, but I really do appreciate you putting them all together into one place.",
    "1718781": "Thank you! At this time ... only one section I update LB - score :) No new ideas which influence on score (in my opinion).",
    "1722007": "I literally started this competition with this post. \nThanks for your hard work, to help people.",
    "1722103": "Great I am happy. Thank you! 👍🙏",
    "1732292": "Thanks for ur nice organizing :)",
    "1732330": "Update in progress :)"
  },
  "source": "meta"
}