{
  "id": 75573,
  "title": "How to deal with texts shown in images",
  "url": "/competitions/humpback-whale-identification/discussion/75573",
  "author_name": "",
  "post_date": "2018-12-23T13:11:32.699389300Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I was thinking to try to remove the text in training set. But I found that there are also some images with texts in test set. Would you suggest to remove the border for training images for training or just keep the texts since there are also texts in test images.</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "444183",
      "postDate": "12/23/2018 13:11:32",
      "content": "<p>I was thinking to try to remove the text in training set. But I found that there are also some images with texts in test set. Would you suggest to remove the border for training images for training or just keep the texts since there are also texts in test images.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "I was thinking to try to remove the text in training set. But I found that there are also some images with texts in test set. Would you suggest to remove the border for training images for training or just keep the texts since there are also texts in test images.\n\nThanks!",
      "votes": null
    },
    {
      "id": "444450",
      "postDate": "12/24/2018 03:47:53",
      "content": "<p>Ideally keep the text; the best algorithm should be able to see the whale and ignore text!</p>",
      "rawMarkdown": "Ideally keep the text; the best algorithm should be able to see the whale and ignore text!",
      "votes": null
    },
    {
      "id": "444472",
      "postDate": "12/24/2018 04:46:58",
      "content": "<p>Text is present, even in ImageNet. The difference between the whale dataset and ImageNet is two orders of magnitude of image data. Therefore, feature engineering will be a significant portion of this challenge. How you extract useful features will be up to you. If you want to detect text and ablate it, that could be a possibility.</p>",
      "rawMarkdown": "Text is present, even in ImageNet. The difference between the whale dataset and ImageNet is two orders of magnitude of image data. Therefore, feature engineering will be a significant portion of this challenge. How you extract useful features will be up to you. If you want to detect text and ablate it, that could be a possibility.",
      "votes": null
    },
    {
      "id": "444483",
      "postDate": "12/24/2018 05:10:02",
      "content": "<p>Perhaps you can add a attention model in your model, if you are using the NN to do this job.</p>",
      "rawMarkdown": "Perhaps you can add a attention model in your model, if you are using the NN to do this job.",
      "votes": null
    },
    {
      "id": "444488",
      "postDate": "12/24/2018 05:34:36",
      "content": "<p>I disagree, cleaning and preparing the data should be the highest priority. I am using the bounding box model from <a href=\"/martinpiotte\">@martinpiotte</a>. It separates out only the tail section in each image. This allows the algorithm to focus on what is most important: the shape and pattern on the whale tail.</p>",
      "rawMarkdown": "I disagree, cleaning and preparing the data should be the highest priority. I am using the bounding box model from @martinpiotte. It separates out only the tail section in each image. This allows the algorithm to focus on what is most important: the shape and pattern on the whale tail.",
      "votes": null
    },
    {
      "id": "444505",
      "postDate": "12/24/2018 06:35:00",
      "content": "<p>It may not very easy to make use of text info, cuz you make need to build a link between the text (such as: #1122) and the corresponding label. </p>",
      "rawMarkdown": "It may not very easy to make use of text info, cuz you make need to build a link between the text (such as: #1122) and the corresponding label.",
      "votes": null
    },
    {
      "id": "444512",
      "postDate": "12/24/2018 06:59:33",
      "content": "<p>I agree with your point. However, the unseen test set may also contains the text information in the images. I am wondering if we only use the bounding box cropping images for training, will it result in a lower score?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "I agree with your point. However, the unseen test set may also contains the text information in the images. I am wondering if we only use the bounding box cropping images for training, will it result in a lower score?\n\nThanks",
      "votes": null
    },
    {
      "id": "444599",
      "postDate": "12/24/2018 11:03:13",
      "content": "<p>Images with text that are proposed to match other images with text will almost invariably be mismatched, thus will lower scores. @Brian is correct that algorithms should focus on the whale only, localization is important</p>",
      "rawMarkdown": "Images with text that are proposed to match other images with text will almost invariably be mismatched, thus will lower scores. @Brian is correct that algorithms should focus on the whale only, localization is important",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 444450,
      "author_name": "tedcheese",
      "author_url": "",
      "post_date": "12/24/2018 03:47:53",
      "content": "<p>Ideally keep the text; the best algorithm should be able to see the whale and ignore text!</p>",
      "votes": null,
      "replies": [
        {
          "id": 444488,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/24/2018 05:34:36",
          "content": "<p>I disagree, cleaning and preparing the data should be the highest priority. I am using the bounding box model from <a href=\"/martinpiotte\">@martinpiotte</a>. It separates out only the tail section in each image. This allows the algorithm to focus on what is most important: the shape and pattern on the whale tail.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444512,
          "author_name": "zaynena",
          "author_url": "",
          "post_date": "12/24/2018 06:59:33",
          "content": "<p>I agree with your point. However, the unseen test set may also contains the text information in the images. I am wondering if we only use the bounding box cropping images for training, will it result in a lower score?</p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444599,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "12/24/2018 11:03:13",
          "content": "<p>Images with text that are proposed to match other images with text will almost invariably be mismatched, thus will lower scores. @Brian is correct that algorithms should focus on the whale only, localization is important</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 444472,
      "author_name": "badtyprr",
      "author_url": "",
      "post_date": "12/24/2018 04:46:58",
      "content": "<p>Text is present, even in ImageNet. The difference between the whale dataset and ImageNet is two orders of magnitude of image data. Therefore, feature engineering will be a significant portion of this challenge. How you extract useful features will be up to you. If you want to detect text and ablate it, that could be a possibility.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 444483,
      "author_name": "markhuang",
      "author_url": "",
      "post_date": "12/24/2018 05:10:02",
      "content": "<p>Perhaps you can add a attention model in your model, if you are using the NN to do this job.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 444505,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "12/24/2018 06:35:00",
      "content": "<p>It may not very easy to make use of text info, cuz you make need to build a link between the text (such as: #1122) and the corresponding label. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "444183": "I was thinking to try to remove the text in training set. But I found that there are also some images with texts in test set. Would you suggest to remove the border for training images for training or just keep the texts since there are also texts in test images.\n\nThanks!",
    "444450": "Ideally keep the text; the best algorithm should be able to see the whale and ignore text!",
    "444472": "Text is present, even in ImageNet. The difference between the whale dataset and ImageNet is two orders of magnitude of image data. Therefore, feature engineering will be a significant portion of this challenge. How you extract useful features will be up to you. If you want to detect text and ablate it, that could be a possibility.",
    "444483": "Perhaps you can add a attention model in your model, if you are using the NN to do this job.",
    "444488": "I disagree, cleaning and preparing the data should be the highest priority. I am using the bounding box model from @martinpiotte. It separates out only the tail section in each image. This allows the algorithm to focus on what is most important: the shape and pattern on the whale tail.",
    "444505": "It may not very easy to make use of text info, cuz you make need to build a link between the text (such as: #1122) and the corresponding label.",
    "444512": "I agree with your point. However, the unseen test set may also contains the text information in the images. I am wondering if we only use the bounding box cropping images for training, will it result in a lower score?\n\nThanks",
    "444599": "Images with text that are proposed to match other images with text will almost invariably be mismatched, thus will lower scores. @Brian is correct that algorithms should focus on the whale only, localization is important"
  },
  "source": "meta"
}