{
  "id": 308026,
  "title": "🐧 Things to know before starting image preprocessing",
  "url": "/competitions/happy-whale-and-dolphin/discussion/308026",
  "author_name": "Andrada",
  "post_date": "2022-02-16T18:58:29.089000",
  "votes": 191,
  "comment_count": 51,
  "views": 0,
  "content": "<p>While browsing through the training/test images (yes, I have looked at 70k+ images), I have realised that the <em>preprocessing</em> part of the images might not be as straight forward as I have thought.</p>\n<p>I am sharing here some findings on the pictures (more info analysis <a href=\"https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance\" target=\"_blank\">in my notebook</a>), in the hope of maybe sharing some approaches on how to properly tackle the <em>noice</em> within the images.</p>\n<ul>\n<li><p><strong>image_size</strong> : the images width and height is very different from one picture to another.</p></li>\n<li><p><strong>night view</strong>: not all the pictures were made during the day - some of them were also caught during the night <em>(convert images to B&amp;W?)</em>. - UPDATE: in comments <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">Peter</a> shows that this is a <code>cv2</code> channel view, not an actual night mode.<br>\n<img src=\"https://i.imgur.com/5IWktBo.png\"></p></li>\n<li><p><strong>multiple individuals</strong>: there are some pictures with 2 or more subjects present in it.<br>\n<img src=\"https://i.imgur.com/uQgnu1h.png\"></p></li>\n<li><p><strong>ladscape</strong>: in some of the images the subject is very close, however in others the landscape is much more predominant, which could impose some issues in identifying subtle characteristics within the individual. In some of them we cannot even see the individual. Also - <em>isn't the landcape a possible leak</em>? Meaning the model would learn the overall landscape pattern instead of the actual individual. However, a very close crop would have little to no detail of the actual subject due to poor quality.<br>\n<img src=\"https://i.imgur.com/Gt9F5nZ.png\"></p></li>\n<li><p><strong>image annotations</strong>: there are some images that have digital marking on them that could pollute the algorithm.<br>\n<img src=\"https://i.imgur.com/tQj1qxf.png\"></p></li>\n<li><p><strong>image duplicates</strong>: there are images where you would think that they are identical copies - in fact, they are not. But they are pictures taken moments appart, so the differences between them are extremely subtle. This could mess up the CV score <em>(erase them?)</em>.<br>\n<img src=\"https://i.imgur.com/psUxHr5.png\"></p></li>\n<li><p><strong>lighting</strong>: I am talking here about the lighting of the surroundings. Sometimes the water can be greenish, sometimes bluish, sometimes pink (because of a sunset for example). Again .. possible fix would be converting to B&amp;W?<br>\n<img src=\"https://i.imgur.com/Lcc3CeQ.png\"></p></li>\n<li><p><strong>🐧penguins!🐧</strong>:  there are penguins in some of the pictures. Adorable!!!<br>\n<img src=\"https://i.imgur.com/dayk0R9.png\"></p></li>\n<li><p><strong>people</strong>: there are also people (tourists and scientists) within the images (not as adorable as the penguins tho).<br>\n<img src=\"https://i.imgur.com/IViKFpo.png\"></p></li>\n<li><p><strong>ice</strong>: initially I thought I would see only water, however, for the individuals that live mostly in arctic waters, many pictures contain ice.<br>\n<img src=\"https://i.imgur.com/MnTjREq.png\"></p></li>\n<li><p><strong>and … this?</strong> 😅<br>\n<img src=\"https://i.imgur.com/8iFjmuw.png\"></p></li>\n</ul>",
  "messages": [
    {
      "id": 1693578,
      "postDate": "2022-02-16T18:58:29.090Z",
      "content": "<p>While browsing through the training/test images (yes, I have looked at 70k+ images), I have realised that the <em>preprocessing</em> part of the images might not be as straight forward as I have thought.</p>\n<p>I am sharing here some findings on the pictures (more info analysis <a href=\"https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance\" target=\"_blank\">in my notebook</a>), in the hope of maybe sharing some approaches on how to properly tackle the <em>noice</em> within the images.</p>\n<ul>\n<li><p><strong>image_size</strong> : the images width and height is very different from one picture to another.</p></li>\n<li><p><strong>night view</strong>: not all the pictures were made during the day - some of them were also caught during the night <em>(convert images to B&amp;W?)</em>. - UPDATE: in comments <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">Peter</a> shows that this is a <code>cv2</code> channel view, not an actual night mode.<br>\n<img src=\"https://i.imgur.com/5IWktBo.png\"></p></li>\n<li><p><strong>multiple individuals</strong>: there are some pictures with 2 or more subjects present in it.<br>\n<img src=\"https://i.imgur.com/uQgnu1h.png\"></p></li>\n<li><p><strong>ladscape</strong>: in some of the images the subject is very close, however in others the landscape is much more predominant, which could impose some issues in identifying subtle characteristics within the individual. In some of them we cannot even see the individual. Also - <em>isn't the landcape a possible leak</em>? Meaning the model would learn the overall landscape pattern instead of the actual individual. However, a very close crop would have little to no detail of the actual subject due to poor quality.<br>\n<img src=\"https://i.imgur.com/Gt9F5nZ.png\"></p></li>\n<li><p><strong>image annotations</strong>: there are some images that have digital marking on them that could pollute the algorithm.<br>\n<img src=\"https://i.imgur.com/tQj1qxf.png\"></p></li>\n<li><p><strong>image duplicates</strong>: there are images where you would think that they are identical copies - in fact, they are not. But they are pictures taken moments appart, so the differences between them are extremely subtle. This could mess up the CV score <em>(erase them?)</em>.<br>\n<img src=\"https://i.imgur.com/psUxHr5.png\"></p></li>\n<li><p><strong>lighting</strong>: I am talking here about the lighting of the surroundings. Sometimes the water can be greenish, sometimes bluish, sometimes pink (because of a sunset for example). Again .. possible fix would be converting to B&amp;W?<br>\n<img src=\"https://i.imgur.com/Lcc3CeQ.png\"></p></li>\n<li><p><strong>🐧penguins!🐧</strong>:  there are penguins in some of the pictures. Adorable!!!<br>\n<img src=\"https://i.imgur.com/dayk0R9.png\"></p></li>\n<li><p><strong>people</strong>: there are also people (tourists and scientists) within the images (not as adorable as the penguins tho).<br>\n<img src=\"https://i.imgur.com/IViKFpo.png\"></p></li>\n<li><p><strong>ice</strong>: initially I thought I would see only water, however, for the individuals that live mostly in arctic waters, many pictures contain ice.<br>\n<img src=\"https://i.imgur.com/MnTjREq.png\"></p></li>\n<li><p><strong>and … this?</strong> 😅<br>\n<img src=\"https://i.imgur.com/8iFjmuw.png\"></p></li>\n</ul>",
      "rawMarkdown": "While browsing through the training/test images (yes, I have looked at 70k+ images), I have realised that the *preprocessing* part of the images might not be as straight forward as I have thought.\n\nI am sharing here some findings on the pictures (more info analysis [in my notebook](https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance)), in the hope of maybe sharing some approaches on how to properly tackle the *noice* within the images.\n\n* **image_size** : the images width and height is very different from one picture to another.\n* **night view**: not all the pictures were made during the day - some of them were also caught during the night *(convert images to B&W?)*. - UPDATE: in comments [Peter](https://www.kaggle.com/pestipeti) shows that this is a `cv2` channel view, not an actual night mode.\n<center><img src=\"https://i.imgur.com/5IWktBo.png\"></center>\n\n* **multiple individuals**: there are some pictures with 2 or more subjects present in it.\n<center><img src=\"https://i.imgur.com/uQgnu1h.png\"></center>\n\n* **ladscape**: in some of the images the subject is very close, however in others the landscape is much more predominant, which could impose some issues in identifying subtle characteristics within the individual. In some of them we cannot even see the individual. Also - *isn't the landcape a possible leak*? Meaning the model would learn the overall landscape pattern instead of the actual individual. However, a very close crop would have little to no detail of the actual subject due to poor quality.\n<center><img src=\"https://i.imgur.com/Gt9F5nZ.png\"></center>\n\n* **image annotations**: there are some images that have digital marking on them that could pollute the algorithm.\n<center><img src=\"https://i.imgur.com/tQj1qxf.png\"></center>\n\n* **image duplicates**: there are images where you would think that they are identical copies - in fact, they are not. But they are pictures taken moments appart, so the differences between them are extremely subtle. This could mess up the CV score *(erase them?)*.\n<center><img src=\"https://i.imgur.com/psUxHr5.png\"></center>\n\n* **lighting**: I am talking here about the lighting of the surroundings. Sometimes the water can be greenish, sometimes bluish, sometimes pink (because of a sunset for example). Again .. possible fix would be converting to B&W?\n<center><img src=\"https://i.imgur.com/Lcc3CeQ.png\"></center>\n\n* **🐧penguins!🐧**:  there are penguins in some of the pictures. Adorable!!!\n<center><img src=\"https://i.imgur.com/dayk0R9.png\"></center>\n\n* **people**: there are also people (tourists and scientists) within the images (not as adorable as the penguins tho).\n<center><img src=\"https://i.imgur.com/IViKFpo.png\"></center>\n\n* **ice**: initially I thought I would see only water, however, for the individuals that live mostly in arctic waters, many pictures contain ice.\n<center><img src=\"https://i.imgur.com/MnTjREq.png\"></center>\n\n* **and ... this?** 😅\n<center><img src=\"https://i.imgur.com/8iFjmuw.png\"></center>",
      "votes": 189
    },
    {
      "id": 1694261,
      "postDate": "2022-02-17T09:52:50.980Z",
      "content": "<p><strong>Chris Deotte talks in <a href=\"https://www.youtube.com/watch?v=XXmujwhjyIo\" target=\"_blank\">this video</a> about the impact of background in classifying an image</strong> - if you have images of cats sitting on cars and dogs sitting on boats, then train a model, you get a good accuracy. </p>\n<p>However, when you test on an image with just a cat (no background), or a cat on a boat, will the model perform as good? I feel like the landscape is a big one to tackle.</p>",
      "rawMarkdown": "**Chris Deotte talks in [this video](https://www.youtube.com/watch?v=XXmujwhjyIo) about the impact of background in classifying an image** - if you have images of cats sitting on cars and dogs sitting on boats, then train a model, you get a good accuracy. \n\nHowever, when you test on an image with just a cat (no background), or a cat on a boat, will the model perform as good? I feel like the landscape is a big one to tackle.",
      "votes": 9,
      "replies": [
        {
          "id": 1695280,
          "postDate": "2022-02-18T04:08:11.517Z",
          "content": "<p>Training a model with some data and then using it for prediction with different-iid data is not recommended. </p>\n<p>However, convolutions are supposed to look and find relevant features for classification, which, in your example, should disregard backgrounds efficiently. Or are they not capable of…?</p>\n<p>For example, just recently I was trying to train from scratch multiple different NNs over CIFAR10 data, and I consistently got low scores. Thus, I think CIFAR10 data is not that iid, because maybe the net does not see enough iid data to clearly extract some key patterns for classification.</p>",
          "rawMarkdown": "Training a model with some data and then using it for prediction with different-iid data is not recommended. \n\nHowever, convolutions are supposed to look and find relevant features for classification, which, in your example, should disregard backgrounds efficiently. Or are they not capable of...?\n\nFor example, just recently I was trying to train from scratch multiple different NNs over CIFAR10 data, and I consistently got low scores. Thus, I think CIFAR10 data is not that iid, because maybe the net does not see enough iid data to clearly extract some key patterns for classification.",
          "votes": 1
        },
        {
          "id": 1695288,
          "postDate": "2022-02-18T04:16:20.157Z",
          "content": "<p><a href=\"https://www.kaggle.com/maulberto3\" target=\"_blank\">@maulberto3</a> I believe the original point in the video was:</p>\n<ul>\n<li>Transformers weren't working well with images</li>\n<li>To investiage, Chris plotted Activation maps onto images</li>\n<li>That made him realise after some experimentation that the problem was with resizing the images </li>\n</ul>\n<blockquote>\n  <p>However, convolutions are supposed to look and find relevant features for classification, which, in your example, should disregard backgrounds efficiently. Or are they not capable of…?</p>\n</blockquote>\n<p>They are capable but need some nudging, also we need to be careful to make sure they don't overfit-which is also when we would see boats becoming a criteria for cats if cats have too many boats in their images. </p>",
          "rawMarkdown": "@maulberto3 I believe the original point in the video was:\n\n- Transformers weren't working well with images\n- To investiage, Chris plotted Activation maps onto images\n- That made him realise after some experimentation that the problem was with resizing the images \n\n> However, convolutions are supposed to look and find relevant features for classification, which, in your example, should disregard backgrounds efficiently. Or are they not capable of…?\n\nThey are capable but need some nudging, also we need to be careful to make sure they don't overfit-which is also when we would see boats becoming a criteria for cats if cats have too many boats in their images. ",
          "votes": 4
        },
        {
          "id": 1696056,
          "postDate": "2022-02-18T15:01:27.687Z",
          "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> Thank you for your explanation. May I also know what is your opinion on my first paragraph?</p>",
          "rawMarkdown": "@init27 Thank you for your explanation. May I also know what is your opinion on my first paragraph?"
        },
        {
          "id": 1697075,
          "postDate": "2022-02-19T10:53:20.783Z",
          "content": "<p><a href=\"https://www.kaggle.com/maulberto3\" target=\"_blank\">@maulberto3</a> I didn't understand what you mean here:</p>\n<blockquote>\n  <p>For example, just recently I was trying to train from scratch multiple different NNs over CIFAR10 data, and I consistently got low scores. Thus, I think CIFAR10 data is not that iid, because maybe the net does not see enough iid data to clearly extract some key patterns for classification.</p>\n</blockquote>\n<p>Were your models not converging? </p>\n<p>From what I understand, iid data can also be mentioned from outside the domain?</p>\n<p>I think what Andrada referred to be either not cropping into train images or also setting up a good TTA for test data. </p>\n<p>But, if we train the model well, it should learn to just identify whales and not stick to the background features as info</p>",
          "rawMarkdown": "@maulberto3 I didn't understand what you mean here:\n\n> For example, just recently I was trying to train from scratch multiple different NNs over CIFAR10 data, and I consistently got low scores. Thus, I think CIFAR10 data is not that iid, because maybe the net does not see enough iid data to clearly extract some key patterns for classification.\n\nWere your models not converging? \n\nFrom what I understand, iid data can also be mentioned from outside the domain?\n\nI think what Andrada referred to be either not cropping into train images or also setting up a good TTA for test data. \n\nBut, if we train the model well, it should learn to just identify whales and not stick to the background features as info",
          "votes": 2
        },
        {
          "id": 1699740,
          "postDate": "2022-02-21T11:57:54.500Z",
          "content": "<p>I'm sorry this might be a newbie question, what's ''iid''?</p>",
          "rawMarkdown": "I'm sorry this might be a newbie question, what's ''iid''?",
          "votes": 1
        },
        {
          "id": 1699751,
          "postDate": "2022-02-21T12:06:52.103Z",
          "content": "<p><strong>I</strong>ndependent and <strong>I</strong>dentically <strong>D</strong>istributed random variables. Read more here: <a href=\"https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables\" target=\"_blank\">Wikipedia Link</a></p>",
          "rawMarkdown": "**I**ndependent and **I**dentically **D**istributed random variables. Read more here: [Wikipedia Link](https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1693681,
      "postDate": "2022-02-16T21:35:03.717Z",
      "content": "<p>it's a messy world, isn't it!? As to that last image… ummmmm… if only we humans were as perfect as machines! 🤕</p>",
      "rawMarkdown": "it's a messy world, isn't it!? As to that last image... ummmmm... if only we humans were as perfect as machines! 🤕",
      "votes": 6,
      "replies": [
        {
          "id": 1693699,
          "postDate": "2022-02-16T22:08:54.150Z",
          "content": "<p>Of, if only the machines were as perfect as us! 😁</p>",
          "rawMarkdown": "Of, if only the machines were as perfect as us! 😁",
          "votes": 4
        },
        {
          "id": 1693703,
          "postDate": "2022-02-16T22:21:20.663Z",
          "content": "<p><a href=\"https://www.kaggle.com/tedcheese\" target=\"_blank\">@tedcheese</a> This is so nice seeing host on discussions. Im sure you are informed by kaggle stuff, but still, try not to accidentally leak something. Again, nice to meet you in here!</p>",
          "rawMarkdown": "@tedcheese This is so nice seeing host on discussions. Im sure you are informed by kaggle stuff, but still, try not to accidentally leak something. Again, nice to meet you in here!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1697875,
      "postDate": "2022-02-19T23:21:16.270Z",
      "content": "<p>In some cases it is better to discard some images, I think</p>",
      "rawMarkdown": "In some cases it is better to discard some images, I think",
      "votes": 1,
      "replies": [
        {
          "id": 1698186,
          "postDate": "2022-02-20T08:00:28.653Z",
          "content": "<p>Yes, I think in the end I will be doing this too.</p>\n<p>For example the image below (image: <code>cd5fe465c60cb9.jpg</code>) is cathegorized as <code>gray_whale</code>, but there is no subject within it. </p>\n<p>Moreover, the subject's <code>individual_id</code> is <code>fc0f7c162cc0</code>, and if we take a look it has 72 more apparitions, so there are planty examples to choose from.</p>\n<p><img src=\"https://i.imgur.com/8iFjmuw.png\"></p>",
          "rawMarkdown": "Yes, I think in the end I will be doing this too.\n\nFor example the image below (image: `cd5fe465c60cb9.jpg`) is cathegorized as `gray_whale`, but there is no subject within it. \n\nMoreover, the subject's `individual_id` is `fc0f7c162cc0`, and if we take a look it has 72 more apparitions, so there are planty examples to choose from.\n\n<img src=\"https://i.imgur.com/8iFjmuw.png\" width=400>",
          "votes": 4
        }
      ]
    },
    {
      "id": 1696870,
      "postDate": "2022-02-19T06:54:29.140Z",
      "content": "<p>Until now, I underestimated the impact of background on the foreground. Great finding andrada</p>",
      "rawMarkdown": "Until now, I underestimated the impact of background on the foreground. Great finding andrada",
      "votes": 1
    },
    {
      "id": 1694716,
      "postDate": "2022-02-17T16:54:39.177Z",
      "content": "<p>How should that clipboard photo be classified? Also Is that mac photo legit? <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>, was this review of training images done manually? If so, wow that's a lot of work.</p>",
      "rawMarkdown": "How should that clipboard photo be classified? Also Is that mac photo legit? @andradaolteanu, was this review of training images done manually? If so, wow that's a lot of work.",
      "votes": 1,
      "replies": [
        {
          "id": 1695598,
          "postDate": "2022-02-18T08:40:05.017Z",
          "content": "<p>Hi Paul!</p>\n<p>Yes, these were done manually. You can see full id and details in <a href=\"https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance/notebook\" target=\"_blank\">my notebook</a> - Chapter 4. Image Analysis, subchapter IV. Other weird image examples. </p>\n<p>My question is - are these examples outliers and should be removed, or are these \"manual errors\" that could be quite frequent and the model should see them? This is what I'm trying to figure out.</p>",
          "rawMarkdown": "Hi Paul!\n\nYes, these were done manually. You can see full id and details in [my notebook](https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance/notebook) - Chapter 4. Image Analysis, subchapter IV. Other weird image examples. \n\nMy question is - are these examples outliers and should be removed, or are these \"manual errors\" that could be quite frequent and the model should see them? This is what I'm trying to figure out.",
          "votes": 1
        },
        {
          "id": 1697042,
          "postDate": "2022-02-19T10:23:27.970Z",
          "content": "<p>Great findings!</p>\n<p>Imho, in large datasets like these, annotating errors are inevitable and usually also present in the test set (sometimes to a lesser extent if test is better curated). If the portion is not too large, models are robust to these and will still generalize well.</p>",
          "rawMarkdown": "Great findings!\n\nImho, in large datasets like these, annotating errors are inevitable and usually also present in the test set (sometimes to a lesser extent if test is better curated). If the portion is not too large, models are robust to these and will still generalize well.",
          "votes": 3
        },
        {
          "id": 1697054,
          "postDate": "2022-02-19T10:36:48.480Z",
          "content": "<blockquote>\n  <p>If the portion is not too large, models are robust to these and will still generalize well.</p>\n</blockquote>\n<p>To share another resource: I remember coming across this website: <a href=\"https://labelerrors.com\" target=\"_blank\">https://labelerrors.com</a></p>\n<p>It was mentioned in this <a href=\"https://arxiv.org/pdf/2103.14749.pdf\" target=\"_blank\">paper</a> which also shares that ImageNet has (I don't recall the exact number) a good number of mislabelled images. </p>\n<p>Almost all of our models pretrained on ImageNet and still work ok 👌</p>",
          "rawMarkdown": "> If the portion is not too large, models are robust to these and will still generalize well.\n\nTo share another resource: I remember coming across this website: https://labelerrors.com\n\nIt was mentioned in this [paper](https://arxiv.org/pdf/2103.14749.pdf) which also shares that ImageNet has (I don't recall the exact number) a good number of mislabelled images. \n\nAlmost all of our models pretrained on ImageNet and still work ok 👌",
          "votes": 1
        }
      ]
    },
    {
      "id": 1693744,
      "postDate": "2022-02-16T22:36:55.597Z",
      "content": "<p>Interesting samples, may I ask how you were able to deduce that they are 'night view'?</p>",
      "rawMarkdown": "Interesting samples, may I ask how you were able to deduce that they are 'night view'?",
      "votes": 1,
      "replies": [
        {
          "id": 1694218,
          "postDate": "2022-02-17T09:06:56.253Z",
          "content": "<p>I <em>think</em> they are night view because they have that classic \"night vision\" aspect you see in movies.</p>\n<p><img src=\"https://i.imgur.com/66iqE6i.png\"></p>",
          "rawMarkdown": "I *think* they are night view because they have that classic \"night vision\" aspect you see in movies.\n\n<img src=\"https://i.imgur.com/66iqE6i.png\" width=400>",
          "votes": 3
        },
        {
          "id": 1694241,
          "postDate": "2022-02-17T09:37:19.447Z",
          "content": "<p>I think cv2 causes this effect. If you read/show a grayscale image without the channel dimension, matplotlib shows in that way.</p>\n<p><img src=\"https://i.imgur.com/Mi0LHlw.jpeg\" alt=\"Imgur\"></p>",
          "rawMarkdown": "I think cv2 causes this effect. If you read/show a grayscale image without the channel dimension, matplotlib shows in that way.\n\n![Imgur](https://i.imgur.com/Mi0LHlw.jpeg)\n",
          "votes": 10
        },
        {
          "id": 1694714,
          "postDate": "2022-02-17T16:53:58.397Z",
          "content": "<p>I think you're right. Infrared cameras only have one channel usually, and this makes sense. Also they might also have a HUD for researchers to use for later reference, which is occurring in some of this B&amp;W, single-channel, infrared photos</p>",
          "rawMarkdown": "I think you're right. Infrared cameras only have one channel usually, and this makes sense. Also they might also have a HUD for researchers to use for later reference, which is occurring in some of this B&W, single-channel, infrared photos",
          "votes": 1
        },
        {
          "id": 1697183,
          "postDate": "2022-02-19T12:38:26.700Z",
          "content": "<p>Peter is right, these are no nightvision images, don't get fooled by matplotlib's false colors! It would be inefficient and dangerous trying to watch whales at night. <br>\nBtw., infrared images would look very different from nightvision. For night shots of land animals, they use infrared lights and sensitive B&amp;W cameras in the near infrared, but the range is very limited and the images noisy and motion-blurred. In passive mid-infrared (no range limitation), you see the heat radiation, all animals would look a bit like belugas (with glowing eyes), but without light reflections, and the contrast might be poor as the skin temperature of sea animals is lower than that of land animals.</p>",
          "rawMarkdown": "Peter is right, these are no nightvision images, don't get fooled by matplotlib's false colors! It would be inefficient and dangerous trying to watch whales at night. \nBtw., infrared images would look very different from nightvision. For night shots of land animals, they use infrared lights and sensitive B&W cameras in the near infrared, but the range is very limited and the images noisy and motion-blurred. In passive mid-infrared (no range limitation), you see the heat radiation, all animals would look a bit like belugas (with glowing eyes), but without light reflections, and the contrast might be poor as the skin temperature of sea animals is lower than that of land animals.",
          "votes": 2
        },
        {
          "id": 1697204,
          "postDate": "2022-02-19T12:52:08.403Z",
          "content": "<p>Thanks for sharing! </p>\n<p>I decided to google a bit more about this. <a href=\"https://www.youtube.com/watch?v=7nsSHuL1NnI\" target=\"_blank\">This</a> 2-minute video tells why IR Cameras are used along with how the IR Imagery looks for animals. </p>\n<p>It's totally different! </p>\n<p>At first just like Andrada, I got tricked into thinking those are Nightvision whale images thanks to my gaming &amp; action movie experience 🤦‍♂️</p>",
          "rawMarkdown": "Thanks for sharing! \n\nI decided to google a bit more about this. [This](https://www.youtube.com/watch?v=7nsSHuL1NnI) 2-minute video tells why IR Cameras are used along with how the IR Imagery looks for animals. \n\nIt's totally different! \n\nAt first just like Andrada, I got tricked into thinking those are Nightvision whale images thanks to my gaming & action movie experience 🤦‍♂️",
          "votes": 1
        },
        {
          "id": 1699742,
          "postDate": "2022-02-21T11:59:31.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> sorry this might be a newbie question, what's \"cv2\"?</p>",
          "rawMarkdown": "@pestipeti sorry this might be a newbie question, what's \"cv2\"?"
        },
        {
          "id": 1699758,
          "postDate": "2022-02-21T12:14:35.200Z",
          "content": "<p>Computer Vision's goto library for working with images: OpenCV. </p>\n<p>By notion, we import it as <code>cv2</code> since that version was most popularised, even though now we're at version <code>3.4.x</code> today :D</p>",
          "rawMarkdown": "Computer Vision's goto library for working with images: OpenCV. \n\nBy notion, we import it as `cv2 ` since that version was most popularised, even though now we're at version `3.4.x` today :D",
          "votes": 1
        },
        {
          "id": 1701039,
          "postDate": "2022-02-22T13:34:40.387Z",
          "content": "<p><a href=\"https://www.kaggle.com/a6893676\" target=\"_blank\">@a6893676</a> cv2 is an image processing library. refer to here <a href=\"https://opencv.org/\" target=\"_blank\">https://opencv.org/</a></p>",
          "rawMarkdown": "@a6893676 cv2 is an image processing library. refer to here https://opencv.org/"
        }
      ]
    },
    {
      "id": 1749983,
      "postDate": "2022-04-09T07:31:06.803Z",
      "content": "<p>Thank you for proposing so many notes that I have never noting. It's a good job.</p>",
      "rawMarkdown": "Thank you for proposing so many notes that I have never noting. It's a good job."
    },
    {
      "id": 1749666,
      "postDate": "2022-04-08T20:44:16.820Z",
      "content": "<p>Insightful!</p>",
      "rawMarkdown": "Insightful!"
    },
    {
      "id": 1731059,
      "postDate": "2022-03-22T00:47:00.743Z",
      "content": "<p>Knowing the data and keeping mind in some of the challenges that we can face is really helpful. Thank you so much for sharing.</p>",
      "rawMarkdown": "Knowing the data and keeping mind in some of the challenges that we can face is really helpful. Thank you so much for sharing."
    },
    {
      "id": 1712488,
      "postDate": "2022-03-05T01:31:50.877Z",
      "content": "<p>Thanks for the detailed explanation.</p>",
      "rawMarkdown": "Thanks for the detailed explanation.",
      "replies": [
        {
          "id": 1735272,
          "postDate": "2022-03-26T04:18:13.817Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1735273,
          "postDate": "2022-03-26T04:19:57.530Z",
          "content": "<p><a href=\"https://www.kaggle.com/weifengbzzz\" target=\"_blank\">@weifengbzzz</a> , 你们队还需要队友吗，我现在single fold 0.773，使用kaggle tensorflow TPU😂</p>",
          "rawMarkdown": "@weifengbzzz , 你们队还需要队友吗，我现在single fold 0.773，使用kaggle tensorflow TPU😂"
        }
      ]
    },
    {
      "id": 1703732,
      "postDate": "2022-02-24T19:06:48.747Z",
      "content": "<p>Useful to get started with image preprocessing</p>",
      "rawMarkdown": "Useful to get started with image preprocessing"
    },
    {
      "id": 1700765,
      "postDate": "2022-02-22T08:26:45.767Z",
      "content": "<p>Very nice investigation about the difficulties and anomalies in this data set (especially the last one). I'm astonished, that the top score is already above 80% for this data.</p>",
      "rawMarkdown": "Very nice investigation about the difficulties and anomalies in this data set (especially the last one). I'm astonished, that the top score is already above 80% for this data."
    },
    {
      "id": 1699854,
      "postDate": "2022-02-21T13:41:05.043Z",
      "content": "<p>Thank you for this informative post <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> </p>",
      "rawMarkdown": "Thank you for this informative post @andradaolteanu "
    },
    {
      "id": 1696136,
      "postDate": "2022-02-18T16:04:04.883Z",
      "content": "<p>Wow the last image is pretty bad 😂</p>\n<p>I had a quick scan of some of the data and found an image with a blurred hand (image: 0091c6a22c29da.jpg).</p>",
      "rawMarkdown": "Wow the last image is pretty bad 😂\n\nI had a quick scan of some of the data and found an image with a blurred hand (image: 0091c6a22c29da.jpg)."
    },
    {
      "id": 1695746,
      "postDate": "2022-02-18T10:41:48.393Z",
      "content": "<p>impressive work <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> thanks for sharing  😊 new follower🙋‍♀️</p>",
      "rawMarkdown": "\nimpressive work @andradaolteanu thanks for sharing  😊 new follower🙋‍♀️"
    },
    {
      "id": 1694910,
      "postDate": "2022-02-17T20:19:45.113Z",
      "content": "<p>Quite Informative and Interesting. <br>\nThanks a lot for sharing.</p>",
      "rawMarkdown": "Quite Informative and Interesting. \nThanks a lot for sharing."
    },
    {
      "id": 1694153,
      "postDate": "2022-02-17T07:39:40.530Z",
      "content": "<p>Good insights, <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>! </p>",
      "rawMarkdown": "Good insights, @andradaolteanu! "
    },
    {
      "id": 1693964,
      "postDate": "2022-02-17T04:57:51.333Z",
      "content": "<p>very well articulated, please write about how to handle noise in the images in your next post. </p>",
      "rawMarkdown": "very well articulated, please write about how to handle noise in the images in your next post. "
    },
    {
      "id": 1693814,
      "postDate": "2022-02-17T01:01:13.067Z",
      "content": "<p>penguins are indeed so adorable 😳😳 thanks for sharing!</p>",
      "rawMarkdown": "penguins are indeed so adorable 😳😳 thanks for sharing!"
    },
    {
      "id": 1700555,
      "postDate": "2022-02-22T05:08:21.110Z",
      "content": "<p>should scaling the same image to multiple resolutions also be a factor while doing preprocessing?</p>",
      "rawMarkdown": "should scaling the same image to multiple resolutions also be a factor while doing preprocessing?",
      "isDeleted": true
    },
    {
      "id": 1695382,
      "postDate": "2022-02-18T05:39:17.447Z",
      "content": "<p>Thank you for sharing those quite interesting information.<br>\nI never suspected that something like the last image would be mixed in the dataset.</p>",
      "rawMarkdown": "Thank you for sharing those quite interesting information.\nI never suspected that something like the last image would be mixed in the dataset.",
      "isDeleted": true
    },
    {
      "id": 1694410,
      "postDate": "2022-02-17T12:12:55.600Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1759837,
      "postDate": "2022-04-19T00:35:28.967Z",
      "content": "<p>Thanks! image_size helps me a lot</p>",
      "rawMarkdown": "Thanks! image_size helps me a lot"
    },
    {
      "id": 1749681,
      "postDate": "2022-04-08T21:29:31.553Z",
      "content": "<p>Thank you for the detailed topic. </p>",
      "rawMarkdown": "Thank you for the detailed topic. "
    },
    {
      "id": 1749581,
      "postDate": "2022-04-08T18:22:00.880Z",
      "content": "<p>Thank you for the information!!</p>",
      "rawMarkdown": "Thank you for the information!!"
    },
    {
      "id": 1707478,
      "postDate": "2022-02-28T13:53:58.370Z",
      "content": "<p>thank you for your hard work!!</p>",
      "rawMarkdown": "thank you for your hard work!!"
    },
    {
      "id": 1702444,
      "postDate": "2022-02-23T15:50:47.060Z",
      "content": "<p>Intresting, thanks for sharing !</p>",
      "rawMarkdown": "Intresting, thanks for sharing !"
    },
    {
      "id": 1702355,
      "postDate": "2022-02-23T14:41:14.917Z",
      "content": "<p>Thank you for the detailed explanation</p>",
      "rawMarkdown": "Thank you for the detailed explanation"
    },
    {
      "id": 1695989,
      "postDate": "2022-02-18T14:24:59.503Z",
      "content": "<p>Thanks for sharing this great finding!</p>",
      "rawMarkdown": "Thanks for sharing this great finding!"
    }
  ],
  "comments": [
    {
      "id": 1694261,
      "author_name": "Andrada",
      "author_url": "",
      "post_date": "2022-02-17T09:52:50.980000",
      "content": "<p><strong>Chris Deotte talks in <a href=\"https://www.youtube.com/watch?v=XXmujwhjyIo\" target=\"_blank\">this video</a> about the impact of background in classifying an image</strong> - if you have images of cats sitting on cars and dogs sitting on boats, then train a model, you get a good accuracy. </p>\n<p>However, when you test on an image with just a cat (no background), or a cat on a boat, will the model perform as good? I feel like the landscape is a big one to tackle.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1695280,
          "author_name": "Mauricio Maroto",
          "author_url": "",
          "post_date": "2022-02-18T04:08:11.517000",
          "content": "<p>Training a model with some data and then using it for prediction with different-iid data is not recommended. </p>\n<p>However, convolutions are supposed to look and find relevant features for classification, which, in your example, should disregard backgrounds efficiently. Or are they not capable of…?</p>\n<p>For example, just recently I was trying to train from scratch multiple different NNs over CIFAR10 data, and I consistently got low scores. Thus, I think CIFAR10 data is not that iid, because maybe the net does not see enough iid data to clearly extract some key patterns for classification.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1695288,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-18T04:16:20.157000",
          "content": "<p><a href=\"https://www.kaggle.com/maulberto3\" target=\"_blank\">@maulberto3</a> I believe the original point in the video was:</p>\n<ul>\n<li>Transformers weren't working well with images</li>\n<li>To investiage, Chris plotted Activation maps onto images</li>\n<li>That made him realise after some experimentation that the problem was with resizing the images </li>\n</ul>\n<blockquote>\n  <p>However, convolutions are supposed to look and find relevant features for classification, which, in your example, should disregard backgrounds efficiently. Or are they not capable of…?</p>\n</blockquote>\n<p>They are capable but need some nudging, also we need to be careful to make sure they don't overfit-which is also when we would see boats becoming a criteria for cats if cats have too many boats in their images. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1696056,
          "author_name": "Mauricio Maroto",
          "author_url": "",
          "post_date": "2022-02-18T15:01:27.687000",
          "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> Thank you for your explanation. May I also know what is your opinion on my first paragraph?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1697075,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-19T10:53:20.783000",
          "content": "<p><a href=\"https://www.kaggle.com/maulberto3\" target=\"_blank\">@maulberto3</a> I didn't understand what you mean here:</p>\n<blockquote>\n  <p>For example, just recently I was trying to train from scratch multiple different NNs over CIFAR10 data, and I consistently got low scores. Thus, I think CIFAR10 data is not that iid, because maybe the net does not see enough iid data to clearly extract some key patterns for classification.</p>\n</blockquote>\n<p>Were your models not converging? </p>\n<p>From what I understand, iid data can also be mentioned from outside the domain?</p>\n<p>I think what Andrada referred to be either not cropping into train images or also setting up a good TTA for test data. </p>\n<p>But, if we train the model well, it should learn to just identify whales and not stick to the background features as info</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1699740,
          "author_name": "Peter CXL",
          "author_url": "",
          "post_date": "2022-02-21T11:57:54.500000",
          "content": "<p>I'm sorry this might be a newbie question, what's ''iid''?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1699751,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-21T12:06:52.103000",
          "content": "<p><strong>I</strong>ndependent and <strong>I</strong>dentically <strong>D</strong>istributed random variables. Read more here: <a href=\"https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables\" target=\"_blank\">Wikipedia Link</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1693681,
      "author_name": "Ted Cheeseman",
      "author_url": "",
      "post_date": "2022-02-16T21:35:03.717000",
      "content": "<p>it's a messy world, isn't it!? As to that last image… ummmmm… if only we humans were as perfect as machines! 🤕</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1693699,
          "author_name": "Andrada",
          "author_url": "",
          "post_date": "2022-02-16T22:08:54.150000",
          "content": "<p>Of, if only the machines were as perfect as us! 😁</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1693703,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2022-02-16T22:21:20.663000",
          "content": "<p><a href=\"https://www.kaggle.com/tedcheese\" target=\"_blank\">@tedcheese</a> This is so nice seeing host on discussions. Im sure you are informed by kaggle stuff, but still, try not to accidentally leak something. Again, nice to meet you in here!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1697875,
      "author_name": "Andres Giraldo",
      "author_url": "",
      "post_date": "2022-02-19T23:21:16.270000",
      "content": "<p>In some cases it is better to discard some images, I think</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1698186,
          "author_name": "Andrada",
          "author_url": "",
          "post_date": "2022-02-20T08:00:28.653000",
          "content": "<p>Yes, I think in the end I will be doing this too.</p>\n<p>For example the image below (image: <code>cd5fe465c60cb9.jpg</code>) is cathegorized as <code>gray_whale</code>, but there is no subject within it. </p>\n<p>Moreover, the subject's <code>individual_id</code> is <code>fc0f7c162cc0</code>, and if we take a look it has 72 more apparitions, so there are planty examples to choose from.</p>\n<p><img src=\"https://i.imgur.com/8iFjmuw.png\"></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1696870,
      "author_name": "Muhammad Sameer Ali Khan",
      "author_url": "",
      "post_date": "2022-02-19T06:54:29.140000",
      "content": "<p>Until now, I underestimated the impact of background on the foreground. Great finding andrada</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1694716,
      "author_name": "Paul",
      "author_url": "",
      "post_date": "2022-02-17T16:54:39.177000",
      "content": "<p>How should that clipboard photo be classified? Also Is that mac photo legit? <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>, was this review of training images done manually? If so, wow that's a lot of work.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1695598,
          "author_name": "Andrada",
          "author_url": "",
          "post_date": "2022-02-18T08:40:05.017000",
          "content": "<p>Hi Paul!</p>\n<p>Yes, these were done manually. You can see full id and details in <a href=\"https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance/notebook\" target=\"_blank\">my notebook</a> - Chapter 4. Image Analysis, subchapter IV. Other weird image examples. </p>\n<p>My question is - are these examples outliers and should be removed, or are these \"manual errors\" that could be quite frequent and the model should see them? This is what I'm trying to figure out.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1697042,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2022-02-19T10:23:27.970000",
          "content": "<p>Great findings!</p>\n<p>Imho, in large datasets like these, annotating errors are inevitable and usually also present in the test set (sometimes to a lesser extent if test is better curated). If the portion is not too large, models are robust to these and will still generalize well.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1697054,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-19T10:36:48.480000",
          "content": "<blockquote>\n  <p>If the portion is not too large, models are robust to these and will still generalize well.</p>\n</blockquote>\n<p>To share another resource: I remember coming across this website: <a href=\"https://labelerrors.com\" target=\"_blank\">https://labelerrors.com</a></p>\n<p>It was mentioned in this <a href=\"https://arxiv.org/pdf/2103.14749.pdf\" target=\"_blank\">paper</a> which also shares that ImageNet has (I don't recall the exact number) a good number of mislabelled images. </p>\n<p>Almost all of our models pretrained on ImageNet and still work ok 👌</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1693744,
      "author_name": "datta",
      "author_url": "",
      "post_date": "2022-02-16T22:36:55.597000",
      "content": "<p>Interesting samples, may I ask how you were able to deduce that they are 'night view'?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1694218,
          "author_name": "Andrada",
          "author_url": "",
          "post_date": "2022-02-17T09:06:56.253000",
          "content": "<p>I <em>think</em> they are night view because they have that classic \"night vision\" aspect you see in movies.</p>\n<p><img src=\"https://i.imgur.com/66iqE6i.png\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1694241,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2022-02-17T09:37:19.447000",
          "content": "<p>I think cv2 causes this effect. If you read/show a grayscale image without the channel dimension, matplotlib shows in that way.</p>\n<p><img src=\"https://i.imgur.com/Mi0LHlw.jpeg\" alt=\"Imgur\"></p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1694714,
          "author_name": "Paul",
          "author_url": "",
          "post_date": "2022-02-17T16:53:58.397000",
          "content": "<p>I think you're right. Infrared cameras only have one channel usually, and this makes sense. Also they might also have a HUD for researchers to use for later reference, which is occurring in some of this B&amp;W, single-channel, infrared photos</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1697183,
          "author_name": "Marius Wanko",
          "author_url": "",
          "post_date": "2022-02-19T12:38:26.700000",
          "content": "<p>Peter is right, these are no nightvision images, don't get fooled by matplotlib's false colors! It would be inefficient and dangerous trying to watch whales at night. <br>\nBtw., infrared images would look very different from nightvision. For night shots of land animals, they use infrared lights and sensitive B&amp;W cameras in the near infrared, but the range is very limited and the images noisy and motion-blurred. In passive mid-infrared (no range limitation), you see the heat radiation, all animals would look a bit like belugas (with glowing eyes), but without light reflections, and the contrast might be poor as the skin temperature of sea animals is lower than that of land animals.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1697204,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-19T12:52:08.403000",
          "content": "<p>Thanks for sharing! </p>\n<p>I decided to google a bit more about this. <a href=\"https://www.youtube.com/watch?v=7nsSHuL1NnI\" target=\"_blank\">This</a> 2-minute video tells why IR Cameras are used along with how the IR Imagery looks for animals. </p>\n<p>It's totally different! </p>\n<p>At first just like Andrada, I got tricked into thinking those are Nightvision whale images thanks to my gaming &amp; action movie experience 🤦‍♂️</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1699742,
          "author_name": "Peter CXL",
          "author_url": "",
          "post_date": "2022-02-21T11:59:31.163000",
          "content": "<p><a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> sorry this might be a newbie question, what's \"cv2\"?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1699758,
          "author_name": "Sanyam Bhutani",
          "author_url": "",
          "post_date": "2022-02-21T12:14:35.200000",
          "content": "<p>Computer Vision's goto library for working with images: OpenCV. </p>\n<p>By notion, we import it as <code>cv2</code> since that version was most popularised, even though now we're at version <code>3.4.x</code> today :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1701039,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2022-02-22T13:34:40.387000",
          "content": "<p><a href=\"https://www.kaggle.com/a6893676\" target=\"_blank\">@a6893676</a> cv2 is an image processing library. refer to here <a href=\"https://opencv.org/\" target=\"_blank\">https://opencv.org/</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1749983,
      "author_name": "JCH2021",
      "author_url": "",
      "post_date": "2022-04-09T07:31:06.803000",
      "content": "<p>Thank you for proposing so many notes that I have never noting. It's a good job.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1749666,
      "author_name": "Aditya Jain",
      "author_url": "",
      "post_date": "2022-04-08T20:44:16.820000",
      "content": "<p>Insightful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1731059,
      "author_name": "Niketan Moon",
      "author_url": "",
      "post_date": "2022-03-22T00:47:00.743000",
      "content": "<p>Knowing the data and keeping mind in some of the challenges that we can face is really helpful. Thank you so much for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1712488,
      "author_name": "WeiFengBZzz",
      "author_url": "",
      "post_date": "2022-03-05T01:31:50.877000",
      "content": "<p>Thanks for the detailed explanation.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1735272,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-03-26T04:18:13.817000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1735273,
          "author_name": "huyahuya",
          "author_url": "",
          "post_date": "2022-03-26T04:19:57.530000",
          "content": "<p><a href=\"https://www.kaggle.com/weifengbzzz\" target=\"_blank\">@weifengbzzz</a> , 你们队还需要队友吗，我现在single fold 0.773，使用kaggle tensorflow TPU😂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1703732,
      "author_name": "SANKET SABOO",
      "author_url": "",
      "post_date": "2022-02-24T19:06:48.747000",
      "content": "<p>Useful to get started with image preprocessing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1700765,
      "author_name": "Gordon Bee",
      "author_url": "",
      "post_date": "2022-02-22T08:26:45.767000",
      "content": "<p>Very nice investigation about the difficulties and anomalies in this data set (especially the last one). I'm astonished, that the top score is already above 80% for this data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1699854,
      "author_name": "Fozan",
      "author_url": "",
      "post_date": "2022-02-21T13:41:05.043000",
      "content": "<p>Thank you for this informative post <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1696136,
      "author_name": "goobill",
      "author_url": "",
      "post_date": "2022-02-18T16:04:04.883000",
      "content": "<p>Wow the last image is pretty bad 😂</p>\n<p>I had a quick scan of some of the data and found an image with a blurred hand (image: 0091c6a22c29da.jpg).</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1695746,
      "author_name": "Aruna S",
      "author_url": "",
      "post_date": "2022-02-18T10:41:48.393000",
      "content": "<p>impressive work <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> thanks for sharing  😊 new follower🙋‍♀️</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1694910,
      "author_name": "Pratinav Seth",
      "author_url": "",
      "post_date": "2022-02-17T20:19:45.113000",
      "content": "<p>Quite Informative and Interesting. <br>\nThanks a lot for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1694153,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2022-02-17T07:39:40.530000",
      "content": "<p>Good insights, <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1693964,
      "author_name": "Ananya Pandey",
      "author_url": "",
      "post_date": "2022-02-17T04:57:51.333000",
      "content": "<p>very well articulated, please write about how to handle noise in the images in your next post. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1693814,
      "author_name": "WONJUNY",
      "author_url": "",
      "post_date": "2022-02-17T01:01:13.067000",
      "content": "<p>penguins are indeed so adorable 😳😳 thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1700555,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-22T05:08:21.110000",
      "content": "<p>should scaling the same image to multiple resolutions also be a factor while doing preprocessing?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1695382,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-18T05:39:17.447000",
      "content": "<p>Thank you for sharing those quite interesting information.<br>\nI never suspected that something like the last image would be mixed in the dataset.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1694410,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-02-17T12:12:55.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1759837,
      "author_name": "蒲钦",
      "author_url": "",
      "post_date": "2022-04-19T00:35:28.967000",
      "content": "<p>Thanks! image_size helps me a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1749681,
      "author_name": "Muhammed Tausif",
      "author_url": "",
      "post_date": "2022-04-08T21:29:31.553000",
      "content": "<p>Thank you for the detailed topic. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1749581,
      "author_name": "Abhinav Kumar",
      "author_url": "",
      "post_date": "2022-04-08T18:22:00.880000",
      "content": "<p>Thank you for the information!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1707478,
      "author_name": "hara",
      "author_url": "",
      "post_date": "2022-02-28T13:53:58.370000",
      "content": "<p>thank you for your hard work!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1702444,
      "author_name": "Amine SNOUSSI",
      "author_url": "",
      "post_date": "2022-02-23T15:50:47.060000",
      "content": "<p>Intresting, thanks for sharing !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1702355,
      "author_name": "HIRO",
      "author_url": "",
      "post_date": "2022-02-23T14:41:14.917000",
      "content": "<p>Thank you for the detailed explanation</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1695989,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2022-02-18T14:24:59.503000",
      "content": "<p>Thanks for sharing this great finding!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1693578": "While browsing through the training/test images (yes, I have looked at 70k+ images), I have realised that the *preprocessing* part of the images might not be as straight forward as I have thought.\n\nI am sharing here some findings on the pictures (more info analysis [in my notebook](https://www.kaggle.com/andradaolteanu/whales-dolphins-effnet-embedding-cos-distance)), in the hope of maybe sharing some approaches on how to properly tackle the *noice* within the images.\n\n* **image_size** : the images width and height is very different from one picture to another.\n* **night view**: not all the pictures were made during the day - some of them were also caught during the night *(convert images to B&W?)*. - UPDATE: in comments [Peter](https://www.kaggle.com/pestipeti) shows that this is a `cv2` channel view, not an actual night mode.\n<center><img src=\"https://i.imgur.com/5IWktBo.png\"></center>\n\n* **multiple individuals**: there are some pictures with 2 or more subjects present in it.\n<center><img src=\"https://i.imgur.com/uQgnu1h.png\"></center>\n\n* **ladscape**: in some of the images the subject is very close, however in others the landscape is much more predominant, which could impose some issues in identifying subtle characteristics within the individual. In some of them we cannot even see the individual. Also - *isn't the landcape a possible leak*? Meaning the model would learn the overall landscape pattern instead of the actual individual. However, a very close crop would have little to no detail of the actual subject due to poor quality.\n<center><img src=\"https://i.imgur.com/Gt9F5nZ.png\"></center>\n\n* **image annotations**: there are some images that have digital marking on them that could pollute the algorithm.\n<center><img src=\"https://i.imgur.com/tQj1qxf.png\"></center>\n\n* **image duplicates**: there are images where you would think that they are identical copies - in fact, they are not. But they are pictures taken moments appart, so the differences between them are extremely subtle. This could mess up the CV score *(erase them?)*.\n<center><img src=\"https://i.imgur.com/psUxHr5.png\"></center>\n\n* **lighting**: I am talking here about the lighting of the surroundings. Sometimes the water can be greenish, sometimes bluish, sometimes pink (because of a sunset for example). Again .. possible fix would be converting to B&W?\n<center><img src=\"https://i.imgur.com/Lcc3CeQ.png\"></center>\n\n* **🐧penguins!🐧**:  there are penguins in some of the pictures. Adorable!!!\n<center><img src=\"https://i.imgur.com/dayk0R9.png\"></center>\n\n* **people**: there are also people (tourists and scientists) within the images (not as adorable as the penguins tho).\n<center><img src=\"https://i.imgur.com/IViKFpo.png\"></center>\n\n* **ice**: initially I thought I would see only water, however, for the individuals that live mostly in arctic waters, many pictures contain ice.\n<center><img src=\"https://i.imgur.com/MnTjREq.png\"></center>\n\n* **and ... this?** 😅\n<center><img src=\"https://i.imgur.com/8iFjmuw.png\"></center>",
    "1694261": "**Chris Deotte talks in [this video](https://www.youtube.com/watch?v=XXmujwhjyIo) about the impact of background in classifying an image** - if you have images of cats sitting on cars and dogs sitting on boats, then train a model, you get a good accuracy. \n\nHowever, when you test on an image with just a cat (no background), or a cat on a boat, will the model perform as good? I feel like the landscape is a big one to tackle.",
    "1693681": "it's a messy world, isn't it!? As to that last image... ummmmm... if only we humans were as perfect as machines! 🤕",
    "1697875": "In some cases it is better to discard some images, I think",
    "1696870": "Until now, I underestimated the impact of background on the foreground. Great finding andrada",
    "1694716": "How should that clipboard photo be classified? Also Is that mac photo legit? @andradaolteanu, was this review of training images done manually? If so, wow that's a lot of work.",
    "1693744": "Interesting samples, may I ask how you were able to deduce that they are 'night view'?",
    "1749983": "Thank you for proposing so many notes that I have never noting. It's a good job.",
    "1749666": "Insightful!",
    "1731059": "Knowing the data and keeping mind in some of the challenges that we can face is really helpful. Thank you so much for sharing.",
    "1712488": "Thanks for the detailed explanation.",
    "1703732": "Useful to get started with image preprocessing",
    "1700765": "Very nice investigation about the difficulties and anomalies in this data set (especially the last one). I'm astonished, that the top score is already above 80% for this data.",
    "1699854": "Thank you for this informative post @andradaolteanu ",
    "1696136": "Wow the last image is pretty bad 😂\n\nI had a quick scan of some of the data and found an image with a blurred hand (image: 0091c6a22c29da.jpg).",
    "1695746": "\nimpressive work @andradaolteanu thanks for sharing  😊 new follower🙋‍♀️",
    "1694910": "Quite Informative and Interesting. \nThanks a lot for sharing.",
    "1694153": "Good insights, @andradaolteanu! ",
    "1693964": "very well articulated, please write about how to handle noise in the images in your next post. ",
    "1693814": "penguins are indeed so adorable 😳😳 thanks for sharing!",
    "1700555": "should scaling the same image to multiple resolutions also be a factor while doing preprocessing?",
    "1695382": "Thank you for sharing those quite interesting information.\nI never suspected that something like the last image would be mixed in the dataset.",
    "1694410": "",
    "1759837": "Thanks! image_size helps me a lot",
    "1749681": "Thank you for the detailed topic. ",
    "1749581": "Thank you for the information!!",
    "1707478": "thank you for your hard work!!",
    "1702444": "Intresting, thanks for sharing !",
    "1702355": "Thank you for the detailed explanation",
    "1695989": "Thanks for sharing this great finding!"
  }
}