{
  "id": 106038,
  "title": "Need some suggestions ? ",
  "url": "/competitions/kuzushiji-recognition/discussion/106038",
  "author_name": "",
  "post_date": "2019-08-27T20:04:45.886250300Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi,\nI am a beginner in computer vision and object detection. I have some experience in making image classifiers in Keras using transfer learning but object detection is completely new for me. This competition seemed interesting so I started working on it, exploring similar past competitions on Kaggle. </p>\n\n<p>From what I have learned, I think the following steps need to be followed:\n1. Cropping the characters using annotations given, from the train images.\n2. training a CNN based on these cropped images.\n3. Dividing the test images into boxes and using every box as an image for our CNN model.</p>\n\n<p>Is this the approach or am I missing something here ? \nAlso, which object detection models(Yolo, R-CNN ) are good and easy to implement ? \nIf you can share any simple implementation of object detection models on custom datasets like this, that'd be great !</p>",
  "messages": [
    {
      "id": "609526",
      "postDate": "08/27/2019 20:04:45",
      "content": "<p>Hi,\nI am a beginner in computer vision and object detection. I have some experience in making image classifiers in Keras using transfer learning but object detection is completely new for me. This competition seemed interesting so I started working on it, exploring similar past competitions on Kaggle. </p>\n\n<p>From what I have learned, I think the following steps need to be followed:\n1. Cropping the characters using annotations given, from the train images.\n2. training a CNN based on these cropped images.\n3. Dividing the test images into boxes and using every box as an image for our CNN model.</p>\n\n<p>Is this the approach or am I missing something here ? \nAlso, which object detection models(Yolo, R-CNN ) are good and easy to implement ? \nIf you can share any simple implementation of object detection models on custom datasets like this, that'd be great !</p>",
      "rawMarkdown": "Hi,\nI am a beginner in computer vision and object detection. I have some experience in making image classifiers in Keras using transfer learning but object detection is completely new for me. This competition seemed interesting so I started working on it, exploring similar past competitions on Kaggle. \n\nFrom what I have learned, I think the following steps need to be followed:\n1. Cropping the characters using annotations given, from the train images.\n2. training a CNN based on these cropped images.\n3. Dividing the test images into boxes and using every box as an image for our CNN model.\n\nIs this the approach or am I missing something here ? \nAlso, which object detection models(Yolo, R-CNN ) are good and easy to implement ? \nIf you can share any simple implementation of object detection models on custom datasets like this, that'd be great !",
      "votes": null
    },
    {
      "id": "609834",
      "postDate": "08/28/2019 06:29:15",
      "content": "<p>Hi,</p>\n\n<p>The approach you are taking is the old approach where a <strong>sliding window</strong> was used to determine the location of the object.</p>\n\n<p>You should use the newer approaches like YOLO and R-CNN where the CNN works like a Sliding window.\nThe network output contains both the location and classification of the object in a structure implemented in CNN Dimensions.</p>\n\n<p>For example, In YOLO's Last layer <em>(depends on implementation)</em> which is a <strong>(7x7x30)</strong> CNN layer, the 7X7 means 49 grid cells. <strong>Each grid cell has 30 values</strong> from CNN's 3rd Dimension. The 30 values have <em>structured meaning,</em> where values form a structure to <em>predict object bounding box size, location and object classifications</em>.</p>\n\n<p>Hope this Helps</p>",
      "rawMarkdown": "Hi,\n\nThe approach you are taking is the old approach where a **sliding window** was used to determine the location of the object.\n\nYou should use the newer approaches like YOLO and R-CNN where the CNN works like a Sliding window.\nThe network output contains both the location and classification of the object in a structure implemented in CNN Dimensions.\n \nFor example, In YOLO's Last layer *(depends on implementation)* which is a **(7x7x30)** CNN layer, the 7X7 means 49 grid cells. **Each grid cell has 30 values** from CNN's 3rd Dimension. The 30 values have *structured meaning,* where values form a structure to *predict object bounding box size, location and object classifications*.\n\nHope this Helps",
      "votes": null
    },
    {
      "id": "609996",
      "postDate": "08/28/2019 09:59:40",
      "content": "<p>But YOLO is trained on some predefined classes. How do we train it on custom dataset ?\nDo we have to create a txt file for every image in the dataset, and then write the class and location of the objects into it ?</p>",
      "rawMarkdown": "But YOLO is trained on some predefined classes. How do we train it on custom dataset ?\nDo we have to create a txt file for every image in the dataset, and then write the class and location of the objects into it ?",
      "votes": null
    },
    {
      "id": "610237",
      "postDate": "08/28/2019 15:10:51",
      "content": "<p>Yes, If you want to <strong>train on custom dataset</strong>, you have to give the model with the <strong>location of every training object and the image</strong> to the model. \nhere is an <strong>example dataset</strong> for detection: <a href=\"https://github.com/moskewcz/VOCdevkit/tree/master/VOC2007\">https://github.com/moskewcz/VOCdevkit/tree/master/VOC2007</a></p>\n\n<p>You can use <strong>DarkFlow implementation</strong>(<a href=\"https://github.com/thtrieu/darkflow\">https://github.com/thtrieu/darkflow</a>) of YOLO to train with this format.\nYou can also use an Annotation tool to annotate and get the data in the right format.</p>",
      "rawMarkdown": "Yes, If you want to **train on custom dataset**, you have to give the model with the **location of every training object and the image** to the model. \nhere is an **example dataset** for detection: https://github.com/moskewcz/VOCdevkit/tree/master/VOC2007\n\nYou can use **DarkFlow implementation**(https://github.com/thtrieu/darkflow) of YOLO to train with this format.\nYou can also use an Annotation tool to annotate and get the data in the right format.",
      "votes": null
    },
    {
      "id": "611478",
      "postDate": "08/29/2019 09:27:18",
      "content": "<p>Thanks a lot for these resources.</p>",
      "rawMarkdown": "Thanks a lot for these resources.",
      "votes": null
    },
    {
      "id": "613014",
      "postDate": "08/30/2019 06:55:44",
      "content": "<p>me too.</p>",
      "rawMarkdown": "me too.",
      "votes": null
    },
    {
      "id": "629908",
      "postDate": "09/19/2019 13:01:40",
      "content": "<p>Excuse me, so the input to the model will be the whole image, and the labels and bboxes for every character?? the same form of train.csv?</p>",
      "rawMarkdown": "Excuse me, so the input to the model will be the whole image, and the labels and bboxes for every character?? the same form of train.csv?",
      "votes": null
    },
    {
      "id": "630176",
      "postDate": "09/19/2019 20:44:48",
      "content": "<p>Yes, one possible approach is to treat this as object detection, then the inputs and outputs are exactly as you described.</p>",
      "rawMarkdown": "Yes, one possible approach is to treat this as object detection, then the inputs and outputs are exactly as you described.",
      "votes": null
    },
    {
      "id": "630754",
      "postDate": "09/20/2019 17:10:05",
      "content": "<p><a href=\"/hassanalsamahi\">@hassanalsamahi</a>  The input to the model is a <strong>resized image</strong> (size depends on model) and output is <strong>labels and object box coordinates prediction</strong>(for models like YOLO) for a set max number of objects.</p>\n\n<p>For <strong>training</strong> object detection models you may have to give <strong>both label and object box coordinates</strong> to the model. so that model can learn from the data.</p>",
      "rawMarkdown": "hassanalsamahi  The input to the model is a **resized image** (size depends on model) and output is **labels and object box coordinates prediction**(for models like YOLO) for a set max number of objects.\n\nFor **training** object detection models you may have to give **both label and object box coordinates** to the model. so that model can learn from the data.",
      "votes": null
    },
    {
      "id": "630780",
      "postDate": "09/20/2019 17:57:39",
      "content": "<p>Okay great, thanks for your explanation</p>",
      "rawMarkdown": "Okay great, thanks for your explanation",
      "votes": null
    },
    {
      "id": "630788",
      "postDate": "09/20/2019 18:06:36",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a>  thanks for your help</p>",
      "rawMarkdown": "lopuhin  thanks for your help",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 609834,
      "author_name": "rameshkamath",
      "author_url": "",
      "post_date": "08/28/2019 06:29:15",
      "content": "<p>Hi,</p>\n\n<p>The approach you are taking is the old approach where a <strong>sliding window</strong> was used to determine the location of the object.</p>\n\n<p>You should use the newer approaches like YOLO and R-CNN where the CNN works like a Sliding window.\nThe network output contains both the location and classification of the object in a structure implemented in CNN Dimensions.</p>\n\n<p>For example, In YOLO's Last layer <em>(depends on implementation)</em> which is a <strong>(7x7x30)</strong> CNN layer, the 7X7 means 49 grid cells. <strong>Each grid cell has 30 values</strong> from CNN's 3rd Dimension. The 30 values have <em>structured meaning,</em> where values form a structure to <em>predict object bounding box size, location and object classifications</em>.</p>\n\n<p>Hope this Helps</p>",
      "votes": null,
      "replies": [
        {
          "id": 609996,
          "author_name": "harshitt21",
          "author_url": "",
          "post_date": "08/28/2019 09:59:40",
          "content": "<p>But YOLO is trained on some predefined classes. How do we train it on custom dataset ?\nDo we have to create a txt file for every image in the dataset, and then write the class and location of the objects into it ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 610237,
          "author_name": "rameshkamath",
          "author_url": "",
          "post_date": "08/28/2019 15:10:51",
          "content": "<p>Yes, If you want to <strong>train on custom dataset</strong>, you have to give the model with the <strong>location of every training object and the image</strong> to the model. \nhere is an <strong>example dataset</strong> for detection: <a href=\"https://github.com/moskewcz/VOCdevkit/tree/master/VOC2007\">https://github.com/moskewcz/VOCdevkit/tree/master/VOC2007</a></p>\n\n<p>You can use <strong>DarkFlow implementation</strong>(<a href=\"https://github.com/thtrieu/darkflow\">https://github.com/thtrieu/darkflow</a>) of YOLO to train with this format.\nYou can also use an Annotation tool to annotate and get the data in the right format.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 611478,
          "author_name": "harshitt21",
          "author_url": "",
          "post_date": "08/29/2019 09:27:18",
          "content": "<p>Thanks a lot for these resources.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 629908,
          "author_name": "hassanalsamahi",
          "author_url": "",
          "post_date": "09/19/2019 13:01:40",
          "content": "<p>Excuse me, so the input to the model will be the whole image, and the labels and bboxes for every character?? the same form of train.csv?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 630176,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "09/19/2019 20:44:48",
          "content": "<p>Yes, one possible approach is to treat this as object detection, then the inputs and outputs are exactly as you described.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 630754,
          "author_name": "rameshkamath",
          "author_url": "",
          "post_date": "09/20/2019 17:10:05",
          "content": "<p><a href=\"/hassanalsamahi\">@hassanalsamahi</a>  The input to the model is a <strong>resized image</strong> (size depends on model) and output is <strong>labels and object box coordinates prediction</strong>(for models like YOLO) for a set max number of objects.</p>\n\n<p>For <strong>training</strong> object detection models you may have to give <strong>both label and object box coordinates</strong> to the model. so that model can learn from the data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 630780,
          "author_name": "hassanalsamahi",
          "author_url": "",
          "post_date": "09/20/2019 17:57:39",
          "content": "<p>Okay great, thanks for your explanation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 630788,
          "author_name": "hassanalsamahi",
          "author_url": "",
          "post_date": "09/20/2019 18:06:36",
          "content": "<p><a href=\"/lopuhin\">@lopuhin</a>  thanks for your help</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 613014,
      "author_name": "need4data",
      "author_url": "",
      "post_date": "08/30/2019 06:55:44",
      "content": "<p>me too.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "609526": "Hi,\nI am a beginner in computer vision and object detection. I have some experience in making image classifiers in Keras using transfer learning but object detection is completely new for me. This competition seemed interesting so I started working on it, exploring similar past competitions on Kaggle. \n\nFrom what I have learned, I think the following steps need to be followed:\n1. Cropping the characters using annotations given, from the train images.\n2. training a CNN based on these cropped images.\n3. Dividing the test images into boxes and using every box as an image for our CNN model.\n\nIs this the approach or am I missing something here ? \nAlso, which object detection models(Yolo, R-CNN ) are good and easy to implement ? \nIf you can share any simple implementation of object detection models on custom datasets like this, that'd be great !",
    "609834": "Hi,\n\nThe approach you are taking is the old approach where a **sliding window** was used to determine the location of the object.\n\nYou should use the newer approaches like YOLO and R-CNN where the CNN works like a Sliding window.\nThe network output contains both the location and classification of the object in a structure implemented in CNN Dimensions.\n \nFor example, In YOLO's Last layer *(depends on implementation)* which is a **(7x7x30)** CNN layer, the 7X7 means 49 grid cells. **Each grid cell has 30 values** from CNN's 3rd Dimension. The 30 values have *structured meaning,* where values form a structure to *predict object bounding box size, location and object classifications*.\n\nHope this Helps",
    "609996": "But YOLO is trained on some predefined classes. How do we train it on custom dataset ?\nDo we have to create a txt file for every image in the dataset, and then write the class and location of the objects into it ?",
    "610237": "Yes, If you want to **train on custom dataset**, you have to give the model with the **location of every training object and the image** to the model. \nhere is an **example dataset** for detection: https://github.com/moskewcz/VOCdevkit/tree/master/VOC2007\n\nYou can use **DarkFlow implementation**(https://github.com/thtrieu/darkflow) of YOLO to train with this format.\nYou can also use an Annotation tool to annotate and get the data in the right format.",
    "611478": "Thanks a lot for these resources.",
    "613014": "me too.",
    "629908": "Excuse me, so the input to the model will be the whole image, and the labels and bboxes for every character?? the same form of train.csv?",
    "630176": "Yes, one possible approach is to treat this as object detection, then the inputs and outputs are exactly as you described.",
    "630754": "hassanalsamahi  The input to the model is a **resized image** (size depends on model) and output is **labels and object box coordinates prediction**(for models like YOLO) for a set max number of objects.\n\nFor **training** object detection models you may have to give **both label and object box coordinates** to the model. so that model can learn from the data.",
    "630780": "Okay great, thanks for your explanation",
    "630788": "lopuhin  thanks for your help"
  },
  "source": "meta"
}