{
  "id": 312853,
  "title": "no background dataset vs detic cropped and backfin dataset train/val question",
  "url": "/competitions/happy-whale-and-dolphin/discussion/312853",
  "author_name": "dragon zhang",
  "post_date": "2022-03-14T11:56:55.373000",
  "votes": 12,
  "comment_count": 13,
  "views": 0,
  "content": "<p>using detic cropped dataset and backfin  obtains a pretty good LB 0.73/0.78</p>\n<p>while using no background  FIN dataset same model, so far one fold cv only 0.569, multiple folds resembling at most around 0.6 ( I guess, without model modification).</p>\n<p>The  question is that the training loss/accuracy no much difference, while the FIN val loss/accuracy much worse.</p>\n<p>What causes the problem?  how to overcome it?</p>\n<p>Thank you for your suggestion.</p>",
  "messages": [
    {
      "id": 1722290,
      "postDate": "2022-03-14T11:56:55.373Z",
      "content": "<p>using detic cropped dataset and backfin  obtains a pretty good LB 0.73/0.78</p>\n<p>while using no background  FIN dataset same model, so far one fold cv only 0.569, multiple folds resembling at most around 0.6 ( I guess, without model modification).</p>\n<p>The  question is that the training loss/accuracy no much difference, while the FIN val loss/accuracy much worse.</p>\n<p>What causes the problem?  how to overcome it?</p>\n<p>Thank you for your suggestion.</p>",
      "rawMarkdown": "using detic cropped dataset and backfin  obtains a pretty good LB 0.73/0.78\n\nwhile using no background  FIN dataset same model, so far one fold cv only 0.569, multiple folds resembling at most around 0.6 ( I guess, without model modification).\n\nThe  question is that the training loss/accuracy no much difference, while the FIN val loss/accuracy much worse.\n\nWhat causes the problem?  how to overcome it?\n\nThank you for your suggestion.",
      "votes": 12
    },
    {
      "id": 1724175,
      "postDate": "2022-03-16T04:06:43.310Z",
      "content": "<p>I think using no background image is not better because some information will be lost (possible)</p>",
      "rawMarkdown": "I think using no background image is not better because some information will be lost (possible)",
      "votes": 1,
      "replies": [
        {
          "id": 1727834,
          "postDate": "2022-03-18T11:16:50.123Z",
          "content": "<p>indeed, there could be a leak between train and test sets if the images are, for example, randomly split, without any consideration on the day the photo was taken. If two photos of an animal are taken the same day, the context (water color or even the land) may be important to identify the animal.</p>",
          "rawMarkdown": "indeed, there could be a leak between train and test sets if the images are, for example, randomly split, without any consideration on the day the photo was taken. If two photos of an animal are taken the same day, the context (water color or even the land) may be important to identify the animal.\n\n ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1736029,
      "postDate": "2022-03-26T20:53:20.363Z",
      "content": "<p>I'm not sure if you're still wondering about this, but here's a thought. In many cases, the identifying feature of a dorsal fin is its trailing edge. For example, dolphins that have many distinctive notches in on the edge of their fin are easily identifiable. The popular <a href=\"https://onlinelibrary.wiley.com/doi/epdf/10.1111/mms.12849\" target=\"_blank\">FinFindR algorithm</a> is built on this principle. The datasets with the background removed may have subtly muted this edge information. </p>",
      "rawMarkdown": "I'm not sure if you're still wondering about this, but here's a thought. In many cases, the identifying feature of a dorsal fin is its trailing edge. For example, dolphins that have many distinctive notches in on the edge of their fin are easily identifiable. The popular [FinFindR algorithm](https://onlinelibrary.wiley.com/doi/epdf/10.1111/mms.12849) is built on this principle. The datasets with the background removed may have subtly muted this edge information. ",
      "votes": 2,
      "replies": [
        {
          "id": 1736033,
          "postDate": "2022-03-26T20:59:41.080Z",
          "content": "<p>Very Good Point!</p>",
          "rawMarkdown": "Very Good Point!"
        }
      ]
    },
    {
      "id": 1723496,
      "postDate": "2022-03-15T13:18:12.127Z",
      "content": "<p><strong>Two questions:</strong></p>\n<ol>\n<li><p>Am I correct in interpreting the info<br>\n  detic cropped dataset      LB 0.73<br>\n  backfin                               LB 0.78</p></li>\n<li><p>Does backfin dataset contain all the images ?</p></li>\n</ol>",
      "rawMarkdown": "**Two questions:**\n\n1. Am I correct in interpreting the info\n      detic cropped dataset      LB 0.73\n      backfin                               LB 0.78\n\n2. Does backfin dataset contain all the images ?"
    },
    {
      "id": 1722469,
      "postDate": "2022-03-14T15:11:11.627Z",
      "content": "<p>I am using pytorch.<br>\nSo… first of all, for me detic is nowhere that good at any case, not even my better detic or the new crowd source is that good, fincrop data is way way better, the difference is like 100-150 points which is huge!, I tried to do fin dataset on tensorflow, it did not go that well, but it was still a bit better than detic, I dont know what you guys are doing to  pull that score with detic, fincrop should obviously be a better way of generalization, however one can argue that detic or full cut provides more information (it covers more area of whale/dolphin).</p>",
      "rawMarkdown": "I am using pytorch.\nSo... first of all, for me detic is nowhere that good at any case, not even my better detic or the new crowd source is that good, fincrop data is way way better, the difference is like 100-150 points which is huge!, I tried to do fin dataset on tensorflow, it did not go that well, but it was still a bit better than detic, I dont know what you guys are doing to  pull that score with detic, fincrop should obviously be a better way of generalization, however one can argue that detic or full cut provides more information (it covers more area of whale/dolphin).",
      "replies": [
        {
          "id": 1722502,
          "postDate": "2022-03-14T15:36:45.410Z",
          "content": "<p><a href=\"https://www.kaggle.com/adnanpen/background-removed-happywhale-dataset\" target=\"_blank\">https://www.kaggle.com/adnanpen/background-removed-happywhale-dataset</a></p>\n<p>what I posted is not well stated.   the dataset FIN  referring to the tfrecords format dataset converted from above removal background dataset. the LB I got so far is worse than detic cropped dataset. <br>\n my best LB is from Backfin dataset.</p>",
          "rawMarkdown": "https://www.kaggle.com/adnanpen/background-removed-happywhale-dataset\n\nwhat I posted is not well stated.   the dataset FIN  referring to the tfrecords format dataset converted from above removal background dataset. the LB I got so far is worse than detic cropped dataset. \n my best LB is from Backfin dataset."
        },
        {
          "id": 1722510,
          "postDate": "2022-03-14T15:46:23.323Z",
          "content": "<p>yeah, same here, after removing the background made the CV way worse in my case too, which was definitely not expected… because logically ocean does nothing but add bias to our model.</p>",
          "rawMarkdown": "yeah, same here, after removing the background made the CV way worse in my case too, which was definitely not expected... because logically ocean does nothing but add bias to our model.",
          "votes": 1
        },
        {
          "id": 1722548,
          "postDate": "2022-03-14T16:28:15.610Z",
          "content": "<p>my guess is that the cropped no background image edges are not reasonably cut.</p>",
          "rawMarkdown": "my guess is that the cropped no background image edges are not reasonably cut."
        },
        {
          "id": 1722574,
          "postDate": "2022-03-14T17:04:34.663Z",
          "content": "<p>Ofc, there is plenty of data which is bad there, but most of them are workable, especially considering that there should also be a boost from non-ocean, I did try mixing no-background images with normal images, it did improve the score but there was no boost, it was just the score which we would get anyways if we did not include no-background images</p>",
          "rawMarkdown": "Ofc, there is plenty of data which is bad there, but most of them are workable, especially considering that there should also be a boost from non-ocean, I did try mixing no-background images with normal images, it did improve the score but there was no boost, it was just the score which we would get anyways if we did not include no-background images"
        },
        {
          "id": 1726050,
          "postDate": "2022-03-17T15:57:12.417Z",
          "content": "<p>I have two bets. I think the nobg dataset confuses the model as they are all resized to the maximum size (while in real life, they could be small or large) and have a white background, so unless you do a similar technique to the test dataset, you won't get good results.</p>",
          "rawMarkdown": "I have two bets. I think the nobg dataset confuses the model as they are all resized to the maximum size (while in real life, they could be small or large) and have a white background, so unless you do a similar technique to the test dataset, you won't get good results.",
          "votes": 1
        },
        {
          "id": 1726182,
          "postDate": "2022-03-17T18:05:18.313Z",
          "content": "<p>I would like to suggest that we are already changing our test set too accordingly, you would not want to do such a huge change and train a model and then run on different images.</p>",
          "rawMarkdown": "I would like to suggest that we are already changing our test set too accordingly, you would not want to do such a huge change and train a model and then run on different images.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1724322,
      "postDate": "2022-03-16T07:25:01.837Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1724175,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2022-03-16T04:06:43.310000",
      "content": "<p>I think using no background image is not better because some information will be lost (possible)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1727834,
          "author_name": "FabienDaniel",
          "author_url": "",
          "post_date": "2022-03-18T11:16:50.123000",
          "content": "<p>indeed, there could be a leak between train and test sets if the images are, for example, randomly split, without any consideration on the day the photo was taken. If two photos of an animal are taken the same day, the context (water color or even the land) may be important to identify the animal.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1736029,
      "author_name": "Phil Patton",
      "author_url": "",
      "post_date": "2022-03-26T20:53:20.363000",
      "content": "<p>I'm not sure if you're still wondering about this, but here's a thought. In many cases, the identifying feature of a dorsal fin is its trailing edge. For example, dolphins that have many distinctive notches in on the edge of their fin are easily identifiable. The popular <a href=\"https://onlinelibrary.wiley.com/doi/epdf/10.1111/mms.12849\" target=\"_blank\">FinFindR algorithm</a> is built on this principle. The datasets with the background removed may have subtly muted this edge information. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1736033,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-26T20:59:41.080000",
          "content": "<p>Very Good Point!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1723496,
      "author_name": "Balaji Selvaraj",
      "author_url": "",
      "post_date": "2022-03-15T13:18:12.127000",
      "content": "<p><strong>Two questions:</strong></p>\n<ol>\n<li><p>Am I correct in interpreting the info<br>\n  detic cropped dataset      LB 0.73<br>\n  backfin                               LB 0.78</p></li>\n<li><p>Does backfin dataset contain all the images ?</p></li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1722469,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-03-14T15:11:11.627000",
      "content": "<p>I am using pytorch.<br>\nSo… first of all, for me detic is nowhere that good at any case, not even my better detic or the new crowd source is that good, fincrop data is way way better, the difference is like 100-150 points which is huge!, I tried to do fin dataset on tensorflow, it did not go that well, but it was still a bit better than detic, I dont know what you guys are doing to  pull that score with detic, fincrop should obviously be a better way of generalization, however one can argue that detic or full cut provides more information (it covers more area of whale/dolphin).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1722502,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-03-14T15:36:45.410000",
          "content": "<p><a href=\"https://www.kaggle.com/adnanpen/background-removed-happywhale-dataset\" target=\"_blank\">https://www.kaggle.com/adnanpen/background-removed-happywhale-dataset</a></p>\n<p>what I posted is not well stated.   the dataset FIN  referring to the tfrecords format dataset converted from above removal background dataset. the LB I got so far is worse than detic cropped dataset. <br>\n my best LB is from Backfin dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722510,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-14T15:46:23.323000",
          "content": "<p>yeah, same here, after removing the background made the CV way worse in my case too, which was definitely not expected… because logically ocean does nothing but add bias to our model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1722548,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-03-14T16:28:15.610000",
          "content": "<p>my guess is that the cropped no background image edges are not reasonably cut.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1722574,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-14T17:04:34.663000",
          "content": "<p>Ofc, there is plenty of data which is bad there, but most of them are workable, especially considering that there should also be a boost from non-ocean, I did try mixing no-background images with normal images, it did improve the score but there was no boost, it was just the score which we would get anyways if we did not include no-background images</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1726050,
          "author_name": "Paul",
          "author_url": "",
          "post_date": "2022-03-17T15:57:12.417000",
          "content": "<p>I have two bets. I think the nobg dataset confuses the model as they are all resized to the maximum size (while in real life, they could be small or large) and have a white background, so unless you do a similar technique to the test dataset, you won't get good results.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1726182,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2022-03-17T18:05:18.313000",
          "content": "<p>I would like to suggest that we are already changing our test set too accordingly, you would not want to do such a huge change and train a model and then run on different images.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1724322,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-16T07:25:01.837000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1722290": "using detic cropped dataset and backfin  obtains a pretty good LB 0.73/0.78\n\nwhile using no background  FIN dataset same model, so far one fold cv only 0.569, multiple folds resembling at most around 0.6 ( I guess, without model modification).\n\nThe  question is that the training loss/accuracy no much difference, while the FIN val loss/accuracy much worse.\n\nWhat causes the problem?  how to overcome it?\n\nThank you for your suggestion.",
    "1724175": "I think using no background image is not better because some information will be lost (possible)",
    "1736029": "I'm not sure if you're still wondering about this, but here's a thought. In many cases, the identifying feature of a dorsal fin is its trailing edge. For example, dolphins that have many distinctive notches in on the edge of their fin are easily identifiable. The popular [FinFindR algorithm](https://onlinelibrary.wiley.com/doi/epdf/10.1111/mms.12849) is built on this principle. The datasets with the background removed may have subtly muted this edge information. ",
    "1723496": "**Two questions:**\n\n1. Am I correct in interpreting the info\n      detic cropped dataset      LB 0.73\n      backfin                               LB 0.78\n\n2. Does backfin dataset contain all the images ?",
    "1722469": "I am using pytorch.\nSo... first of all, for me detic is nowhere that good at any case, not even my better detic or the new crowd source is that good, fincrop data is way way better, the difference is like 100-150 points which is huge!, I tried to do fin dataset on tensorflow, it did not go that well, but it was still a bit better than detic, I dont know what you guys are doing to  pull that score with detic, fincrop should obviously be a better way of generalization, however one can argue that detic or full cut provides more information (it covers more area of whale/dolphin).",
    "1724322": ""
  }
}