{
  "id": 313987,
  "title": "Main Booster - Dataset",
  "url": "/competitions/happy-whale-and-dolphin/discussion/313987",
  "author_name": "",
  "post_date": "2022-03-20T08:25:19.270175500Z",
  "votes": 32,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Best performance is achieved by choosing the proper dataset.</p>\n<p>Detic &gt; yolo if you are using the following dataset<br>\n<a href=\"https://www.kaggle.com/code/lextoumbourou/happywhale-tfrecords-with-bounding-boxes/notebook\" target=\"_blank\">https://www.kaggle.com/code/lextoumbourou/happywhale-tfrecords-with-bounding-boxes/notebook</a><br>\nLB - 0.747</p>\n<p>It takes minor effort to <br>\nuse the bounding box info from  <a href=\"https://www.kaggle.com/jpbremer/fullbodywhaleannotations\" target=\"_blank\">https://www.kaggle.com/jpbremer/fullbodywhaleannotations</a><br>\nLB - 0.786</p>\n<p>During the analysis, I used exactly the same training environment and I changed only the dataset.<br>\nThe performance is observed in a single model EffnetB7 (600 image-size) trained for 40 epochs. Single fold(45k images) </p>\n<p>There seem to be a few better datasets that are the key boosters {that the top kernels are using}.<br>\nFrom your experiments, which dataset is yielding better results for you?</p>",
  "messages": [
    {
      "id": "1729574",
      "postDate": "03/20/2022 08:25:19",
      "content": "<p>Best performance is achieved by choosing the proper dataset.</p>\n<p>Detic &gt; yolo if you are using the following dataset<br>\n<a href=\"https://www.kaggle.com/code/lextoumbourou/happywhale-tfrecords-with-bounding-boxes/notebook\" target=\"_blank\">https://www.kaggle.com/code/lextoumbourou/happywhale-tfrecords-with-bounding-boxes/notebook</a><br>\nLB - 0.747</p>\n<p>It takes minor effort to <br>\nuse the bounding box info from  <a href=\"https://www.kaggle.com/jpbremer/fullbodywhaleannotations\" target=\"_blank\">https://www.kaggle.com/jpbremer/fullbodywhaleannotations</a><br>\nLB - 0.786</p>\n<p>During the analysis, I used exactly the same training environment and I changed only the dataset.<br>\nThe performance is observed in a single model EffnetB7 (600 image-size) trained for 40 epochs. Single fold(45k images) </p>\n<p>There seem to be a few better datasets that are the key boosters {that the top kernels are using}.<br>\nFrom your experiments, which dataset is yielding better results for you?</p>",
      "rawMarkdown": "Best performance is achieved by choosing the proper dataset.\n\nDetic > yolo if you are using the following dataset\nhttps://www.kaggle.com/code/lextoumbourou/happywhale-tfrecords-with-bounding-boxes/notebook\nLB - 0.747\n\nIt takes minor effort to \nuse the bounding box info from  https://www.kaggle.com/jpbremer/fullbodywhaleannotations\nLB - 0.786\n\nDuring the analysis, I used exactly the same training environment and I changed only the dataset.\nThe performance is observed in a single model EffnetB7 (600 image-size) trained for 40 epochs. Single fold(45k images) \n\nThere seem to be a few better datasets that are the key boosters {that the top kernels are using}.\nFrom your experiments, which dataset is yielding better results for you?",
      "votes": null
    },
    {
      "id": "1729606",
      "postDate": "03/20/2022 09:23:20",
      "content": "<p>What do you mean - add dataset to Detic?<br>\nIs this score for one fold?</p>",
      "rawMarkdown": "What do you mean - add dataset to Detic?\nIs this score for one fold?",
      "votes": null
    },
    {
      "id": "1729631",
      "postDate": "03/20/2022 10:01:01",
      "content": "<p>I mean.. In the tf records we have bounding box info of detic and yolo.. U can also add the bounding boxes in <a href=\"https://www.kaggle.com/jpbremer/fullbodywhaleannotations\" target=\"_blank\">https://www.kaggle.com/jpbremer/fullbodywhaleannotations</a> and use it instead of detic Or yolo</p>",
      "rawMarkdown": "I mean.. In the tf records we have bounding box info of detic and yolo.. U can also add the bounding boxes in https://www.kaggle.com/jpbremer/fullbodywhaleannotations and use it instead of detic Or yolo",
      "votes": null
    },
    {
      "id": "1729633",
      "postDate": "03/20/2022 10:02:42",
      "content": "<p>Results mentioned are from 1 fold <br>\n~45k images in training. I used a different splitting strategy. </p>",
      "rawMarkdown": "Results mentioned are from 1 fold \n~45k images in training. I used a different splitting strategy.",
      "votes": null
    },
    {
      "id": "1729636",
      "postDate": "03/20/2022 10:07:43",
      "content": "<p>I am asking because LB - 0.786 score from one fold is very good. Thank you for answer.</p>",
      "rawMarkdown": "I am asking because LB - 0.786 score from one fold is very good. Thank you for answer.",
      "votes": null
    },
    {
      "id": "1729639",
      "postDate": "03/20/2022 10:14:09",
      "content": "<p>I think your dataset will be a game-changer. Eagerly waiting to hear results about it. </p>\n<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> <br>\nIf you don't mind, can you say<br>\nWhat was the best booster in your experiments? <br>\nwhat dataset are you using? </p>",
      "rawMarkdown": "I think your dataset will be a game-changer. Eagerly waiting to hear results about it. \n\n@remekkinas \nIf you don't mind, can you say\nWhat was the best booster in your experiments? \nwhat dataset are you using?",
      "votes": null
    },
    {
      "id": "1729653",
      "postDate": "03/20/2022 10:20:50",
      "content": "<p>We use custom dataset. Not published so far (because it is a result of many tricks we made). It will be published after competition. We have made many experiments with DS and found optimal one. Hovewer I know that it can be better (we have problem to jump over 0.820 - I think this is still DS)  … so I am looking for next step. </p>\n<p>I do not use dataset with removed backgroud yet. </p>\n<p>The best booster? Still the same … work on Dataset. All major jumps in score in our case is connected with DS change.</p>",
      "rawMarkdown": "We use custom dataset. Not published so far (because it is a result of many tricks we made). It will be published after competition. We have made many experiments with DS and found optimal one. Hovewer I know that it can be better (we have problem to jump over 0.820 - I think this is still DS)  ... so I am looking for next step. \n\nI do not use dataset with removed backgroud yet. \n\nThe best booster? Still the same ... work on Dataset. All major jumps in score in our case is connected with DS change.",
      "votes": null
    },
    {
      "id": "1730284",
      "postDate": "03/21/2022 04:58:24",
      "content": "<p>A shortcut is to augment your data (train and test).<br>\ne.g for a false positive test image, you can create multiple crops (location and scale.aspect). score each of the crops<br>\nwhich crop is the one that gives the best score? can I make a detector to detect it (to fasten the process)?</p>\n<p>who says we must only use one crop in a image? we can use multiple crop + whole image, etc</p>",
      "rawMarkdown": "A shortcut is to augment your data (train and test).\ne.g for a false positive test image, you can create multiple crops (location and scale.aspect). score each of the crops\nwhich crop is the one that gives the best score? can I make a detector to detect it (to fasten the process)?\n\nwho says we must only use one crop in a image? we can use multiple crop + whole image, etc",
      "votes": null
    },
    {
      "id": "1730350",
      "postDate": "03/21/2022 07:02:50",
      "content": "<p>My best single fold is 782lb and 848cv. But 5folds only get 791 in lb</p>",
      "rawMarkdown": "My best single fold is 782lb and 848cv. But 5folds only get 791 in lb",
      "votes": null
    },
    {
      "id": "1730374",
      "postDate": "03/21/2022 07:32:00",
      "content": "<p>TF?<br>\nWhat input resolution?</p>",
      "rawMarkdown": "TF?\nWhat input resolution?",
      "votes": null
    },
    {
      "id": "1730400",
      "postDate": "03/21/2022 08:01:23",
      "content": "<p>could you share tfrecords dataset  so that no wasting platform resource?<br>\nthanks.</p>",
      "rawMarkdown": "could you share tfrecords dataset  so that no wasting platform resource?\nthanks.",
      "votes": null
    },
    {
      "id": "1730430",
      "postDate": "03/21/2022 08:43:56",
      "content": "<p>I experienced a big boost when switching from <a href=\"https://www.kaggle.com/phalanx/whale2-cropped-dataset\" target=\"_blank\">Detic</a> to backfins, which I converted here from backfintfrecords: <a href=\"https://www.kaggle.com/code/clemchris/convert-backfintfrecords\" target=\"_blank\">https://www.kaggle.com/code/clemchris/convert-backfintfrecords</a>. It jumped from LB=0.567 (image_size=512, \"tf_efficientnet_b4\", batch_size=16) to LB=0.656 (image_size=384, \"tf_efficientnet_b4\"): <a href=\"https://www.kaggle.com/code/clemchris/pytorch-backfin-convnext-arcface/notebook\" target=\"_blank\">https://www.kaggle.com/code/clemchris/pytorch-backfin-convnext-arcface/notebook</a></p>",
      "rawMarkdown": "I experienced a big boost when switching from [Detic](https://www.kaggle.com/phalanx/whale2-cropped-dataset) to backfins, which I converted here from backfintfrecords: https://www.kaggle.com/code/clemchris/convert-backfintfrecords. It jumped from LB=0.567 (image_size=512, \"tf_efficientnet_b4\", batch_size=16) to LB=0.656 (image_size=384, \"tf_efficientnet_b4\"): https://www.kaggle.com/code/clemchris/pytorch-backfin-convnext-arcface/notebook",
      "votes": null
    },
    {
      "id": "1730733",
      "postDate": "03/21/2022 15:10:31",
      "content": "<p>I just hope no one is touching the test data.<br>\nIf someone does handlabelling anyway, how will kaggle catch it?</p>",
      "rawMarkdown": "I just hope no one is touching the test data.\nIf someone does handlabelling anyway, how will kaggle catch it?",
      "votes": null
    },
    {
      "id": "1730812",
      "postDate": "03/21/2022 16:20:58",
      "content": "<p>It's just so stupid of that person to do that. He isn't going to learn anything. <br>\nSo don't worry about that person. Enjoy learning my friend 👍</p>",
      "rawMarkdown": "It's just so stupid of that person to do that. He isn't going to learn anything. \nSo don't worry about that person. Enjoy learning my friend 👍",
      "votes": null
    },
    {
      "id": "1731900",
      "postDate": "03/22/2022 20:16:17",
      "content": "<p>Hi, how did you manage to handle bounding boxes from here:  <a href=\"https://www.kaggle.com/jpbremer/fullbodywhaleannotations\" target=\"_blank\">https://www.kaggle.com/jpbremer/fullbodywhaleannotations</a>? It seems like it's YOLO outputs, but I get broken crops when I'm trying to crop the image this way.</p>",
      "rawMarkdown": "Hi, how did you manage to handle bounding boxes from here:  https://www.kaggle.com/jpbremer/fullbodywhaleannotations? It seems like it's YOLO outputs, but I get broken crops when I'm trying to crop the image this way.",
      "votes": null
    },
    {
      "id": "1732069",
      "postDate": "03/23/2022 01:13:35",
      "content": "<p>Yes. Input size was 768.</p>",
      "rawMarkdown": "Yes. Input size was 768.",
      "votes": null
    },
    {
      "id": "1732483",
      "postDate": "03/23/2022 12:35:32",
      "content": "<p>Actually its not stupid at all, handlabeling test data can be used for score boost with pseudo, so its like using that data in training, making the generalization better. I experienced a boost of 0.025 points in early stages of pseudo labeling test data, but I am not handlabeling anything, so if someone hand labels predictions which are really bad in test data is a big improvement to both public test and private test because of improved pseudo labels.</p>",
      "rawMarkdown": "Actually its not stupid at all, handlabeling test data can be used for score boost with pseudo, so its like using that data in training, making the generalization better. I experienced a boost of 0.025 points in early stages of pseudo labeling test data, but I am not handlabeling anything, so if someone hand labels predictions which are really bad in test data is a big improvement to both public test and private test because of improved pseudo labels.",
      "votes": null
    },
    {
      "id": "1732492",
      "postDate": "03/23/2022 12:49:06",
      "content": "<p>Pseudo labelling and hand labelling are 2 very different things! In the case of pseudo labelling you label test set samples with your best model, <strong>not by hand</strong>, and then you can include the labelled samples into the training set. <br>\nBut never ever touch the test samples by hand!  </p>",
      "rawMarkdown": "Pseudo labelling and hand labelling are 2 very different things! In the case of pseudo labelling you label test set samples with your best model, **not by hand**, and then you can include the labelled samples into the training set. \nBut never ever touch the test samples by hand!",
      "votes": null
    },
    {
      "id": "1732496",
      "postDate": "03/23/2022 12:56:19",
      "content": "<p>I absolutely agree with you, Allie. Pseudo-labeling is good. I don't like the hand labeling test datasets.</p>",
      "rawMarkdown": "I absolutely agree with you, Allie. Pseudo-labeling is good. I don't like the hand labeling test datasets.",
      "votes": null
    },
    {
      "id": "1732790",
      "postDate": "03/23/2022 18:33:54",
      "content": "<p>I am saying that you can re-make better pseudo with handlabeled changes, its a possibility, I am in no support of this, as I find it unfair.</p>",
      "rawMarkdown": "I am saying that you can re-make better pseudo with handlabeled changes, its a possibility, I am in no support of this, as I find it unfair.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1729606,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "03/20/2022 09:23:20",
      "content": "<p>What do you mean - add dataset to Detic?<br>\nIs this score for one fold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1729631,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/20/2022 10:01:01",
          "content": "<p>I mean.. In the tf records we have bounding box info of detic and yolo.. U can also add the bounding boxes in <a href=\"https://www.kaggle.com/jpbremer/fullbodywhaleannotations\" target=\"_blank\">https://www.kaggle.com/jpbremer/fullbodywhaleannotations</a> and use it instead of detic Or yolo</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729633,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/20/2022 10:02:42",
          "content": "<p>Results mentioned are from 1 fold <br>\n~45k images in training. I used a different splitting strategy. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729636,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/20/2022 10:07:43",
          "content": "<p>I am asking because LB - 0.786 score from one fold is very good. Thank you for answer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729639,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/20/2022 10:14:09",
          "content": "<p>I think your dataset will be a game-changer. Eagerly waiting to hear results about it. </p>\n<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> <br>\nIf you don't mind, can you say<br>\nWhat was the best booster in your experiments? <br>\nwhat dataset are you using? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729653,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/20/2022 10:20:50",
          "content": "<p>We use custom dataset. Not published so far (because it is a result of many tricks we made). It will be published after competition. We have made many experiments with DS and found optimal one. Hovewer I know that it can be better (we have problem to jump over 0.820 - I think this is still DS)  … so I am looking for next step. </p>\n<p>I do not use dataset with removed backgroud yet. </p>\n<p>The best booster? Still the same … work on Dataset. All major jumps in score in our case is connected with DS change.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730350,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "03/21/2022 07:02:50",
          "content": "<p>My best single fold is 782lb and 848cv. But 5folds only get 791 in lb</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1730374,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/21/2022 07:32:00",
          "content": "<p>TF?<br>\nWhat input resolution?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732069,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "03/23/2022 01:13:35",
          "content": "<p>Yes. Input size was 768.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1730284,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/21/2022 04:58:24",
      "content": "<p>A shortcut is to augment your data (train and test).<br>\ne.g for a false positive test image, you can create multiple crops (location and scale.aspect). score each of the crops<br>\nwhich crop is the one that gives the best score? can I make a detector to detect it (to fasten the process)?</p>\n<p>who says we must only use one crop in a image? we can use multiple crop + whole image, etc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1730400,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "03/21/2022 08:01:23",
      "content": "<p>could you share tfrecords dataset  so that no wasting platform resource?<br>\nthanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1730430,
      "author_name": "clemchris",
      "author_url": "",
      "post_date": "03/21/2022 08:43:56",
      "content": "<p>I experienced a big boost when switching from <a href=\"https://www.kaggle.com/phalanx/whale2-cropped-dataset\" target=\"_blank\">Detic</a> to backfins, which I converted here from backfintfrecords: <a href=\"https://www.kaggle.com/code/clemchris/convert-backfintfrecords\" target=\"_blank\">https://www.kaggle.com/code/clemchris/convert-backfintfrecords</a>. It jumped from LB=0.567 (image_size=512, \"tf_efficientnet_b4\", batch_size=16) to LB=0.656 (image_size=384, \"tf_efficientnet_b4\"): <a href=\"https://www.kaggle.com/code/clemchris/pytorch-backfin-convnext-arcface/notebook\" target=\"_blank\">https://www.kaggle.com/code/clemchris/pytorch-backfin-convnext-arcface/notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1730733,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/21/2022 15:10:31",
      "content": "<p>I just hope no one is touching the test data.<br>\nIf someone does handlabelling anyway, how will kaggle catch it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1730812,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/21/2022 16:20:58",
          "content": "<p>It's just so stupid of that person to do that. He isn't going to learn anything. <br>\nSo don't worry about that person. Enjoy learning my friend 👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732483,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/23/2022 12:35:32",
          "content": "<p>Actually its not stupid at all, handlabeling test data can be used for score boost with pseudo, so its like using that data in training, making the generalization better. I experienced a boost of 0.025 points in early stages of pseudo labeling test data, but I am not handlabeling anything, so if someone hand labels predictions which are really bad in test data is a big improvement to both public test and private test because of improved pseudo labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732492,
          "author_name": "blankaf",
          "author_url": "",
          "post_date": "03/23/2022 12:49:06",
          "content": "<p>Pseudo labelling and hand labelling are 2 very different things! In the case of pseudo labelling you label test set samples with your best model, <strong>not by hand</strong>, and then you can include the labelled samples into the training set. <br>\nBut never ever touch the test samples by hand!  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732496,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/23/2022 12:56:19",
          "content": "<p>I absolutely agree with you, Allie. Pseudo-labeling is good. I don't like the hand labeling test datasets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732790,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/23/2022 18:33:54",
          "content": "<p>I am saying that you can re-make better pseudo with handlabeled changes, its a possibility, I am in no support of this, as I find it unfair.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1731900,
      "author_name": "vadimtimakin",
      "author_url": "",
      "post_date": "03/22/2022 20:16:17",
      "content": "<p>Hi, how did you manage to handle bounding boxes from here:  <a href=\"https://www.kaggle.com/jpbremer/fullbodywhaleannotations\" target=\"_blank\">https://www.kaggle.com/jpbremer/fullbodywhaleannotations</a>? It seems like it's YOLO outputs, but I get broken crops when I'm trying to crop the image this way.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1729574": "Best performance is achieved by choosing the proper dataset.\n\nDetic > yolo if you are using the following dataset\nhttps://www.kaggle.com/code/lextoumbourou/happywhale-tfrecords-with-bounding-boxes/notebook\nLB - 0.747\n\nIt takes minor effort to \nuse the bounding box info from  https://www.kaggle.com/jpbremer/fullbodywhaleannotations\nLB - 0.786\n\nDuring the analysis, I used exactly the same training environment and I changed only the dataset.\nThe performance is observed in a single model EffnetB7 (600 image-size) trained for 40 epochs. Single fold(45k images) \n\nThere seem to be a few better datasets that are the key boosters {that the top kernels are using}.\nFrom your experiments, which dataset is yielding better results for you?",
    "1729606": "What do you mean - add dataset to Detic?\nIs this score for one fold?",
    "1729631": "I mean.. In the tf records we have bounding box info of detic and yolo.. U can also add the bounding boxes in https://www.kaggle.com/jpbremer/fullbodywhaleannotations and use it instead of detic Or yolo",
    "1729633": "Results mentioned are from 1 fold \n~45k images in training. I used a different splitting strategy.",
    "1729636": "I am asking because LB - 0.786 score from one fold is very good. Thank you for answer.",
    "1729639": "I think your dataset will be a game-changer. Eagerly waiting to hear results about it. \n\n@remekkinas \nIf you don't mind, can you say\nWhat was the best booster in your experiments? \nwhat dataset are you using?",
    "1729653": "We use custom dataset. Not published so far (because it is a result of many tricks we made). It will be published after competition. We have made many experiments with DS and found optimal one. Hovewer I know that it can be better (we have problem to jump over 0.820 - I think this is still DS)  ... so I am looking for next step. \n\nI do not use dataset with removed backgroud yet. \n\nThe best booster? Still the same ... work on Dataset. All major jumps in score in our case is connected with DS change.",
    "1730284": "A shortcut is to augment your data (train and test).\ne.g for a false positive test image, you can create multiple crops (location and scale.aspect). score each of the crops\nwhich crop is the one that gives the best score? can I make a detector to detect it (to fasten the process)?\n\nwho says we must only use one crop in a image? we can use multiple crop + whole image, etc",
    "1730350": "My best single fold is 782lb and 848cv. But 5folds only get 791 in lb",
    "1730374": "TF?\nWhat input resolution?",
    "1730400": "could you share tfrecords dataset  so that no wasting platform resource?\nthanks.",
    "1730430": "I experienced a big boost when switching from [Detic](https://www.kaggle.com/phalanx/whale2-cropped-dataset) to backfins, which I converted here from backfintfrecords: https://www.kaggle.com/code/clemchris/convert-backfintfrecords. It jumped from LB=0.567 (image_size=512, \"tf_efficientnet_b4\", batch_size=16) to LB=0.656 (image_size=384, \"tf_efficientnet_b4\"): https://www.kaggle.com/code/clemchris/pytorch-backfin-convnext-arcface/notebook",
    "1730733": "I just hope no one is touching the test data.\nIf someone does handlabelling anyway, how will kaggle catch it?",
    "1730812": "It's just so stupid of that person to do that. He isn't going to learn anything. \nSo don't worry about that person. Enjoy learning my friend 👍",
    "1731900": "Hi, how did you manage to handle bounding boxes from here:  https://www.kaggle.com/jpbremer/fullbodywhaleannotations? It seems like it's YOLO outputs, but I get broken crops when I'm trying to crop the image this way.",
    "1732069": "Yes. Input size was 768.",
    "1732483": "Actually its not stupid at all, handlabeling test data can be used for score boost with pseudo, so its like using that data in training, making the generalization better. I experienced a boost of 0.025 points in early stages of pseudo labeling test data, but I am not handlabeling anything, so if someone hand labels predictions which are really bad in test data is a big improvement to both public test and private test because of improved pseudo labels.",
    "1732492": "Pseudo labelling and hand labelling are 2 very different things! In the case of pseudo labelling you label test set samples with your best model, **not by hand**, and then you can include the labelled samples into the training set. \nBut never ever touch the test samples by hand!",
    "1732496": "I absolutely agree with you, Allie. Pseudo-labeling is good. I don't like the hand labeling test datasets.",
    "1732790": "I am saying that you can re-make better pseudo with handlabeled changes, its a possibility, I am in no support of this, as I find it unfair."
  },
  "source": "meta"
}