{
  "id": 309188,
  "title": "EDA in CV ... should be done different way then in tabular data ...",
  "url": "/competitions/happy-whale-and-dolphin/discussion/309188",
  "author_name": "",
  "post_date": "2022-02-22T09:01:47.788478600Z",
  "votes": 17,
  "comment_count": 7,
  "views": 0,
  "content": "<p>EDA in computer vision is different then in tabular data. We must work visually on each step (even visual validation of our model performance). My CV workflow for data:</p>\n<ul>\n<li>look into data - spend 80% working with images (I always dwonload whole dataset and watching on photos - this competition I went through at least 50% of dataset)</li>\n<li>visualize images (or … visualize video - thie was my favourite tool in Tensor Flow competition)</li>\n<li>visualize prediction (classification result, bboxes, segmentations etc.) - score is not enough - we sholud see prediction and well and find anomalies to improve model</li>\n</ul>\n<p>Here I found excellent article about EDA in CV: <a href=\"https://neptune.ai/blog/data-exploration-for-image-segmentation-and-object-detection\" target=\"_blank\">https://neptune.ai/blog/data-exploration-for-image-segmentation-and-object-detection</a></p>\n<p>Two samples of the same ID whale. Any conclusions?</p>\n<p>The same whale ID in dataset but probably you see different whales. When you put them into regonition pipeline … you are in trouble.<br>\n<img src=\"https://i.ibb.co/MkBRXGs/plot-1.jpg\" alt=\"1\"></p>\n<p>The same whale … but two different side for recognition.<br>\n<img src=\"https://i.ibb.co/M2jnwB3/plot.jpg\" alt=\"2\"></p>",
  "messages": [
    {
      "id": "1700791",
      "postDate": "02/22/2022 09:01:47",
      "content": "<p>EDA in computer vision is different then in tabular data. We must work visually on each step (even visual validation of our model performance). My CV workflow for data:</p>\n<ul>\n<li>look into data - spend 80% working with images (I always dwonload whole dataset and watching on photos - this competition I went through at least 50% of dataset)</li>\n<li>visualize images (or … visualize video - thie was my favourite tool in Tensor Flow competition)</li>\n<li>visualize prediction (classification result, bboxes, segmentations etc.) - score is not enough - we sholud see prediction and well and find anomalies to improve model</li>\n</ul>\n<p>Here I found excellent article about EDA in CV: <a href=\"https://neptune.ai/blog/data-exploration-for-image-segmentation-and-object-detection\" target=\"_blank\">https://neptune.ai/blog/data-exploration-for-image-segmentation-and-object-detection</a></p>\n<p>Two samples of the same ID whale. Any conclusions?</p>\n<p>The same whale ID in dataset but probably you see different whales. When you put them into regonition pipeline … you are in trouble.<br>\n<img src=\"https://i.ibb.co/MkBRXGs/plot-1.jpg\" alt=\"1\"></p>\n<p>The same whale … but two different side for recognition.<br>\n<img src=\"https://i.ibb.co/M2jnwB3/plot.jpg\" alt=\"2\"></p>",
      "rawMarkdown": "EDA in computer vision is different then in tabular data. We must work visually on each step (even visual validation of our model performance). My CV workflow for data:\n- look into data - spend 80% working with images (I always dwonload whole dataset and watching on photos - this competition I went through at least 50% of dataset)\n- visualize images (or ... visualize video - thie was my favourite tool in Tensor Flow competition)\n- visualize prediction (classification result, bboxes, segmentations etc.) - score is not enough - we sholud see prediction and well and find anomalies to improve model\n\nHere I found excellent article about EDA in CV: https://neptune.ai/blog/data-exploration-for-image-segmentation-and-object-detection\n \nTwo samples of the same ID whale. Any conclusions?\n\nThe same whale ID in dataset but probably you see different whales. When you put them into regonition pipeline ... you are in trouble.\n![1](https://i.ibb.co/MkBRXGs/plot-1.jpg)\n\nThe same whale ... but two different side for recognition.\n![2](https://i.ibb.co/M2jnwB3/plot.jpg)",
      "votes": null
    },
    {
      "id": "1700833",
      "postDate": "02/22/2022 09:40:33",
      "content": "<p>this is not exactly recognition due to dataset setting. It is more like similarity searching.</p>",
      "rawMarkdown": "this is not exactly recognition due to dataset setting. It is more like similarity searching.",
      "votes": null
    },
    {
      "id": "1700839",
      "postDate": "02/22/2022 09:50:06",
      "content": "<p>and? What can you find simillar whale in test if in training (baseline) one whale ID is represented by different whales?</p>\n<p>Let's say you want to find similar people to .. Andrew NG (in test) …. but training dataset has:</p>\n<ul>\n<li>ID(Ag) -&gt; 10 photos of Lecun</li>\n<li>ID(Ag) -&gt; 1 photo of Andrew Ng</li>\n<li>ID(Ag) -&gt; 20 photos of Dragon Zhang</li>\n<li>ID(Ag) -&gt; 15 photos of Remek Kinas</li>\n</ul>\n<p>Training:</p>\n<ul>\n<li>ID(Ag)-&gt; Andrew Ng -&gt; maybe we find Andrew Ng but it depends on solution </li>\n<li>ID (Remek) --&gt; Andrew Ng</li>\n</ul>",
      "rawMarkdown": "and? What can you find simillar whale in test if in training (baseline) one whale ID is represented by different whales?\n\nLet's say you want to find similar people to .. Andrew NG (in test) .... but training dataset has:\n- ID(Ag) -> 10 photos of Lecun\n- ID(Ag) -> 1 photo of Andrew Ng\n- ID(Ag) -> 20 photos of Dragon Zhang\n- ID(Ag) -> 15 photos of Remek Kinas\n\nTraining:\n- ID(Ag)-> Andrew Ng -> maybe we find Andrew Ng but it depends on solution \n- ID (Remek) --> Andrew Ng",
      "votes": null
    },
    {
      "id": "1701728",
      "postDate": "02/23/2022 02:53:00",
      "content": "<p>Are you saying that some of the data is labelled wrong? different whales with same ID? </p>",
      "rawMarkdown": "Are you saying that some of the data is labelled wrong? different whales with same ID?",
      "votes": null
    },
    {
      "id": "1701999",
      "postDate": "02/23/2022 08:31:04",
      "content": "<p>Probably yes. Today I will look into deeper and show cases. </p>",
      "rawMarkdown": "Probably yes. Today I will look into deeper and show cases.",
      "votes": null
    },
    {
      "id": "1713311",
      "postDate": "03/05/2022 21:47:42",
      "content": "<p>100% agree. Adding a \"plot inference images\" function to debug a CV pipeline is often very important (both observed in professional and Kaggle projects). I also like to quickly write a streamlit app to explore the images dataset, nothing better than understanding how the images look. 👌</p>",
      "rawMarkdown": "100% agree. Adding a \"plot inference images\" function to debug a CV pipeline is often very important (both observed in professional and Kaggle projects). I also like to quickly write a streamlit app to explore the images dataset, nothing better than understanding how the images look. 👌",
      "votes": null
    },
    {
      "id": "1713312",
      "postDate": "03/05/2022 21:50:21",
      "content": "<p>I agree in 100%. In CV looking on images is the best way to validate many things. We as a human are able to see details and patterns …. In tabular data we do not have such abilities so we draw many charts to understand data.</p>",
      "rawMarkdown": "I agree in 100%. In CV looking on images is the best way to validate many things. We as a human are able to see details and patterns .... In tabular data we do not have such abilities so we draw many charts to understand data.",
      "votes": null
    },
    {
      "id": "1713314",
      "postDate": "03/05/2022 22:12:35",
      "content": "<p>That's one reason I prefer computer vision competitions nowadays, I can use my own \"visual system\" to explore the data without a lot of fancy transformations. 😄 </p>",
      "rawMarkdown": "That's one reason I prefer computer vision competitions nowadays, I can use my own \"visual system\" to explore the data without a lot of fancy transformations. 😄",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1700833,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/22/2022 09:40:33",
      "content": "<p>this is not exactly recognition due to dataset setting. It is more like similarity searching.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1700839,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/22/2022 09:50:06",
          "content": "<p>and? What can you find simillar whale in test if in training (baseline) one whale ID is represented by different whales?</p>\n<p>Let's say you want to find similar people to .. Andrew NG (in test) …. but training dataset has:</p>\n<ul>\n<li>ID(Ag) -&gt; 10 photos of Lecun</li>\n<li>ID(Ag) -&gt; 1 photo of Andrew Ng</li>\n<li>ID(Ag) -&gt; 20 photos of Dragon Zhang</li>\n<li>ID(Ag) -&gt; 15 photos of Remek Kinas</li>\n</ul>\n<p>Training:</p>\n<ul>\n<li>ID(Ag)-&gt; Andrew Ng -&gt; maybe we find Andrew Ng but it depends on solution </li>\n<li>ID (Remek) --&gt; Andrew Ng</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1701728,
      "author_name": "ashwincheekati",
      "author_url": "",
      "post_date": "02/23/2022 02:53:00",
      "content": "<p>Are you saying that some of the data is labelled wrong? different whales with same ID? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1701999,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/23/2022 08:31:04",
          "content": "<p>Probably yes. Today I will look into deeper and show cases. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1713311,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/05/2022 21:47:42",
      "content": "<p>100% agree. Adding a \"plot inference images\" function to debug a CV pipeline is often very important (both observed in professional and Kaggle projects). I also like to quickly write a streamlit app to explore the images dataset, nothing better than understanding how the images look. 👌</p>",
      "votes": null,
      "replies": [
        {
          "id": 1713312,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "03/05/2022 21:50:21",
          "content": "<p>I agree in 100%. In CV looking on images is the best way to validate many things. We as a human are able to see details and patterns …. In tabular data we do not have such abilities so we draw many charts to understand data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1713314,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/05/2022 22:12:35",
          "content": "<p>That's one reason I prefer computer vision competitions nowadays, I can use my own \"visual system\" to explore the data without a lot of fancy transformations. 😄 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1700791": "EDA in computer vision is different then in tabular data. We must work visually on each step (even visual validation of our model performance). My CV workflow for data:\n- look into data - spend 80% working with images (I always dwonload whole dataset and watching on photos - this competition I went through at least 50% of dataset)\n- visualize images (or ... visualize video - thie was my favourite tool in Tensor Flow competition)\n- visualize prediction (classification result, bboxes, segmentations etc.) - score is not enough - we sholud see prediction and well and find anomalies to improve model\n\nHere I found excellent article about EDA in CV: https://neptune.ai/blog/data-exploration-for-image-segmentation-and-object-detection\n \nTwo samples of the same ID whale. Any conclusions?\n\nThe same whale ID in dataset but probably you see different whales. When you put them into regonition pipeline ... you are in trouble.\n![1](https://i.ibb.co/MkBRXGs/plot-1.jpg)\n\nThe same whale ... but two different side for recognition.\n![2](https://i.ibb.co/M2jnwB3/plot.jpg)",
    "1700833": "this is not exactly recognition due to dataset setting. It is more like similarity searching.",
    "1700839": "and? What can you find simillar whale in test if in training (baseline) one whale ID is represented by different whales?\n\nLet's say you want to find similar people to .. Andrew NG (in test) .... but training dataset has:\n- ID(Ag) -> 10 photos of Lecun\n- ID(Ag) -> 1 photo of Andrew Ng\n- ID(Ag) -> 20 photos of Dragon Zhang\n- ID(Ag) -> 15 photos of Remek Kinas\n\nTraining:\n- ID(Ag)-> Andrew Ng -> maybe we find Andrew Ng but it depends on solution \n- ID (Remek) --> Andrew Ng",
    "1701728": "Are you saying that some of the data is labelled wrong? different whales with same ID?",
    "1701999": "Probably yes. Today I will look into deeper and show cases.",
    "1713311": "100% agree. Adding a \"plot inference images\" function to debug a CV pipeline is often very important (both observed in professional and Kaggle projects). I also like to quickly write a streamlit app to explore the images dataset, nothing better than understanding how the images look. 👌",
    "1713312": "I agree in 100%. In CV looking on images is the best way to validate many things. We as a human are able to see details and patterns .... In tabular data we do not have such abilities so we draw many charts to understand data.",
    "1713314": "That's one reason I prefer computer vision competitions nowadays, I can use my own \"visual system\" to explore the data without a lot of fancy transformations. 😄"
  },
  "source": "meta"
}