{
  "id": 104124,
  "title": "Which model do you think is better to start with???",
  "url": "/competitions/kuzushiji-recognition/discussion/104124",
  "author_name": "",
  "post_date": "2019-08-14T13:26:17.032754Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello everyone,\nI know many of you have started with segmentation techniques like faster-rcnn and regression based technique like YOLO-v3 or SSD technique, so can you suggest which method will suit better for this problem and explain why/why not??? I am planning to work with YOLO-v3 because I have worked on it for before.</p>",
  "messages": [
    {
      "id": "599086",
      "postDate": "08/14/2019 13:26:17",
      "content": "<p>Hello everyone,\nI know many of you have started with segmentation techniques like faster-rcnn and regression based technique like YOLO-v3 or SSD technique, so can you suggest which method will suit better for this problem and explain why/why not??? I am planning to work with YOLO-v3 because I have worked on it for before.</p>",
      "rawMarkdown": "Hello everyone,\nI know many of you have started with segmentation techniques like faster-rcnn and regression based technique like YOLO-v3 or SSD technique, so can you suggest which method will suit better for this problem and explain why/why not??? I am planning to work with YOLO-v3 because I have worked on it for before.",
      "votes": null
    },
    {
      "id": "599862",
      "postDate": "08/15/2019 12:01:53",
      "content": "<p><a href=\"/shivus\">@shivus</a> do you think this is a object detection problem?</p>",
      "rawMarkdown": "shivus do you think this is a object detection problem?",
      "votes": null
    },
    {
      "id": "599903",
      "postDate": "08/15/2019 12:42:00",
      "content": "<p>YOLO-V3 did not work for me .  The detected bboxes are very bad.  Centernet is a good starter.  You can see @k_mat 's great kernel:  <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">https://www.kaggle.com/kmat2019/centernet-keypoint-detector</a>  .   </p>",
      "rawMarkdown": "YOLO-V3 did not work for me .  The detected bboxes are very bad.  Centernet is a good starter.  You can see @k_mat 's great kernel:  https://www.kaggle.com/kmat2019/centernet-keypoint-detector  .",
      "votes": null
    },
    {
      "id": "599933",
      "postDate": "08/15/2019 13:19:07",
      "content": "<p>Thanks for sharing @huigin. I had a small clarrification. Does this algorithm mentioned the research paper by <a href=\"/anokas\">@anokas</a> work out for this problem?</p>\n\n<ol>\n<li>Train two seperate variational autoencoder on pixel version of KanjiVG and Kuzhushiji-Kanji on 64x64px resolution.</li>\n<li>Train mixture density network to mode P(Znew | Zold) as mixture of gaussians.</li>\n<li>Train sketch RNN to generate Kanji VGG strokes conditioned on either znew or z~new ~P(Znew|Zold).</li>\n</ol>",
      "rawMarkdown": "Thanks for sharing @huigin. I had a small clarrification. Does this algorithm mentioned the research paper by @anokas work out for this problem?\n\n1. Train two seperate variational autoencoder on pixel version of KanjiVG and Kuzhushiji-Kanji on 64x64px resolution.\n2. Train mixture density network to mode P(Znew | Zold) as mixture of gaussians.\n3. Train sketch RNN to generate Kanji VGG strokes conditioned on either znew or z~new ~P(Znew|Zold).",
      "votes": null
    },
    {
      "id": "599935",
      "postDate": "08/15/2019 13:19:29",
      "content": "<p>in research paper: <a href=\"https://arxiv.org/pdf/1812.01718.pdf\">https://arxiv.org/pdf/1812.01718.pdf</a></p>",
      "rawMarkdown": "in research paper: https://arxiv.org/pdf/1812.01718.pdf",
      "votes": null
    },
    {
      "id": "600432",
      "postDate": "08/16/2019 05:46:16",
      "content": "<p>I check this paper, but can not find the bbox detecting  algorithm.  Hence I think it can not work for this problem. </p>",
      "rawMarkdown": "I check this paper, but can not find the bbox detecting  algorithm.  Hence I think it can not work for this problem.",
      "votes": null
    },
    {
      "id": "600947",
      "postDate": "08/16/2019 19:32:40",
      "content": "<p>It is, isn't it? The problem is to localize and classify the characters, which is basically object detection!</p>",
      "rawMarkdown": "It is, isn't it? The problem is to localize and classify the characters, which is basically object detection!",
      "votes": null
    },
    {
      "id": "601389",
      "postDate": "08/17/2019 14:41:13",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a>  thanks for sharing!!!</p>",
      "rawMarkdown": "qinhui1999  thanks for sharing!!!",
      "votes": null
    },
    {
      "id": "601411",
      "postDate": "08/17/2019 15:15:45",
      "content": "<p>Thank you <a href=\"/qinhui1999\">@qinhui1999</a>  for introducing my kernel. </p>\n\n<p>In my opinion, detecting letters is not so difficult, because the size of letters doesn't vary so much in each picture and overlaps between bounding boxes are very small. I rather cared the small object size (more than 200 letters in a picture) than the architecture. </p>\n\n<p>I think any detector works to a certain degree as long as you select an appropriate size of inputs and outputs of the detector for each picture. For instance, if you use YOLO-v3 with input size of 512x512, the maximum detection feature map would be 64x64(stride 8). This detector is not good at detecting the object much smaller than this grid cell. </p>",
      "rawMarkdown": "Thank you @qinhui1999  for introducing my kernel. \n\nIn my opinion, detecting letters is not so difficult, because the size of letters doesn't vary so much in each picture and overlaps between bounding boxes are very small. I rather cared the small object size (more than 200 letters in a picture) than the architecture. \n\nI think any detector works to a certain degree as long as you select an appropriate size of inputs and outputs of the detector for each picture. For instance, if you use YOLO-v3 with input size of 512x512, the maximum detection feature map would be 64x64(stride 8). This detector is not good at detecting the object much smaller than this grid cell.",
      "votes": null
    },
    {
      "id": "601853",
      "postDate": "08/18/2019 08:44:56",
      "content": "<p>Thanks <a href=\"/kmat2019\">@kmat2019</a> 's detailed answer . Yep,  yolo-v3 is not good at detecting such small objects(64*64).  For AP small( &lt; 32*32) and AP medium( &gt;32*32 and &lt;96*96) , <br>\nyolo-v3 got 18.3 and 25.4 scores respective.  <br>\nAnd Centernet got 19.9 and 43 scores respective.<br>\nSo ,if we want  yolo-v3 work, then we should let the feature map bigger than 96*96,  right?</p>",
      "rawMarkdown": "Thanks @kmat2019 's detailed answer . Yep,  yolo-v3 is not good at detecting such small objects(64*64).  For AP small( &lt; 32*32) and AP medium( &gt;32*32 and &lt;96*96) , <br>\nyolo-v3 got 18.3 and 25.4 scores respective.  <br>\nAnd Centernet got 19.9 and 43 scores respective.<br>\nSo ,if we want  yolo-v3 work, then we should let the feature map bigger than 96*96,  right?",
      "votes": null
    },
    {
      "id": "601924",
      "postDate": "08/18/2019 11:15:16",
      "content": "<p>I agree. Controlling the feature map (grid cell) is one possible solution to detect small object. </p>\n\n<p>The other way is zooming in by splitting the picture into several parts. In my kernel, I run the detector approximately 4 times for each picture. If you can know the size of the object in advance, this would be the simplest solution. </p>",
      "rawMarkdown": "I agree. Controlling the feature map (grid cell) is one possible solution to detect small object. \n\nThe other way is zooming in by splitting the picture into several parts. In my kernel, I run the detector approximately 4 times for each picture. If you can know the size of the object in advance, this would be the simplest solution.",
      "votes": null
    },
    {
      "id": "616388",
      "postDate": "09/03/2019 05:08:44",
      "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> <a href=\"/shivus\">@shivus</a>   I think it could work if first use a region proposal network (like R-CNN) first, or have a single pass network like YOLO produce feature vectors in addition to coordinates that are then passed to the autoencoders. Either way, it could definitely work as one head within a multi-objective network.</p>",
      "rawMarkdown": "qinhui1999 @shivus   I think it could work if first use a region proposal network (like R-CNN) first, or have a single pass network like YOLO produce feature vectors in addition to coordinates that are then passed to the autoencoders. Either way, it could definitely work as one head within a multi-objective network.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 599862,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "08/15/2019 12:01:53",
      "content": "<p><a href=\"/shivus\">@shivus</a> do you think this is a object detection problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 600947,
          "author_name": "christianwallenwein",
          "author_url": "",
          "post_date": "08/16/2019 19:32:40",
          "content": "<p>It is, isn't it? The problem is to localize and classify the characters, which is basically object detection!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 599903,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "08/15/2019 12:42:00",
      "content": "<p>YOLO-V3 did not work for me .  The detected bboxes are very bad.  Centernet is a good starter.  You can see @k_mat 's great kernel:  <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">https://www.kaggle.com/kmat2019/centernet-keypoint-detector</a>  .   </p>",
      "votes": null,
      "replies": [
        {
          "id": 599933,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "08/15/2019 13:19:07",
          "content": "<p>Thanks for sharing @huigin. I had a small clarrification. Does this algorithm mentioned the research paper by <a href=\"/anokas\">@anokas</a> work out for this problem?</p>\n\n<ol>\n<li>Train two seperate variational autoencoder on pixel version of KanjiVG and Kuzhushiji-Kanji on 64x64px resolution.</li>\n<li>Train mixture density network to mode P(Znew | Zold) as mixture of gaussians.</li>\n<li>Train sketch RNN to generate Kanji VGG strokes conditioned on either znew or z~new ~P(Znew|Zold).</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 599935,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "08/15/2019 13:19:29",
          "content": "<p>in research paper: <a href=\"https://arxiv.org/pdf/1812.01718.pdf\">https://arxiv.org/pdf/1812.01718.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 600432,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "08/16/2019 05:46:16",
          "content": "<p>I check this paper, but can not find the bbox detecting  algorithm.  Hence I think it can not work for this problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601389,
          "author_name": "shivus",
          "author_url": "",
          "post_date": "08/17/2019 14:41:13",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a>  thanks for sharing!!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 616388,
          "author_name": "deasmhumnha",
          "author_url": "",
          "post_date": "09/03/2019 05:08:44",
          "content": "<p><a href=\"/qinhui1999\">@qinhui1999</a> <a href=\"/shivus\">@shivus</a>   I think it could work if first use a region proposal network (like R-CNN) first, or have a single pass network like YOLO produce feature vectors in addition to coordinates that are then passed to the autoencoders. Either way, it could definitely work as one head within a multi-objective network.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 601411,
      "author_name": "kmat2019",
      "author_url": "",
      "post_date": "08/17/2019 15:15:45",
      "content": "<p>Thank you <a href=\"/qinhui1999\">@qinhui1999</a>  for introducing my kernel. </p>\n\n<p>In my opinion, detecting letters is not so difficult, because the size of letters doesn't vary so much in each picture and overlaps between bounding boxes are very small. I rather cared the small object size (more than 200 letters in a picture) than the architecture. </p>\n\n<p>I think any detector works to a certain degree as long as you select an appropriate size of inputs and outputs of the detector for each picture. For instance, if you use YOLO-v3 with input size of 512x512, the maximum detection feature map would be 64x64(stride 8). This detector is not good at detecting the object much smaller than this grid cell. </p>",
      "votes": null,
      "replies": [
        {
          "id": 601853,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "08/18/2019 08:44:56",
          "content": "<p>Thanks <a href=\"/kmat2019\">@kmat2019</a> 's detailed answer . Yep,  yolo-v3 is not good at detecting such small objects(64*64).  For AP small( &lt; 32*32) and AP medium( &gt;32*32 and &lt;96*96) , <br>\nyolo-v3 got 18.3 and 25.4 scores respective.  <br>\nAnd Centernet got 19.9 and 43 scores respective.<br>\nSo ,if we want  yolo-v3 work, then we should let the feature map bigger than 96*96,  right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 601924,
          "author_name": "kmat2019",
          "author_url": "",
          "post_date": "08/18/2019 11:15:16",
          "content": "<p>I agree. Controlling the feature map (grid cell) is one possible solution to detect small object. </p>\n\n<p>The other way is zooming in by splitting the picture into several parts. In my kernel, I run the detector approximately 4 times for each picture. If you can know the size of the object in advance, this would be the simplest solution. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "599086": "Hello everyone,\nI know many of you have started with segmentation techniques like faster-rcnn and regression based technique like YOLO-v3 or SSD technique, so can you suggest which method will suit better for this problem and explain why/why not??? I am planning to work with YOLO-v3 because I have worked on it for before.",
    "599862": "shivus do you think this is a object detection problem?",
    "599903": "YOLO-V3 did not work for me .  The detected bboxes are very bad.  Centernet is a good starter.  You can see @k_mat 's great kernel:  https://www.kaggle.com/kmat2019/centernet-keypoint-detector  .",
    "599933": "Thanks for sharing @huigin. I had a small clarrification. Does this algorithm mentioned the research paper by @anokas work out for this problem?\n\n1. Train two seperate variational autoencoder on pixel version of KanjiVG and Kuzhushiji-Kanji on 64x64px resolution.\n2. Train mixture density network to mode P(Znew | Zold) as mixture of gaussians.\n3. Train sketch RNN to generate Kanji VGG strokes conditioned on either znew or z~new ~P(Znew|Zold).",
    "599935": "in research paper: https://arxiv.org/pdf/1812.01718.pdf",
    "600432": "I check this paper, but can not find the bbox detecting  algorithm.  Hence I think it can not work for this problem.",
    "600947": "It is, isn't it? The problem is to localize and classify the characters, which is basically object detection!",
    "601389": "qinhui1999  thanks for sharing!!!",
    "601411": "Thank you @qinhui1999  for introducing my kernel. \n\nIn my opinion, detecting letters is not so difficult, because the size of letters doesn't vary so much in each picture and overlaps between bounding boxes are very small. I rather cared the small object size (more than 200 letters in a picture) than the architecture. \n\nI think any detector works to a certain degree as long as you select an appropriate size of inputs and outputs of the detector for each picture. For instance, if you use YOLO-v3 with input size of 512x512, the maximum detection feature map would be 64x64(stride 8). This detector is not good at detecting the object much smaller than this grid cell.",
    "601853": "Thanks @kmat2019 's detailed answer . Yep,  yolo-v3 is not good at detecting such small objects(64*64).  For AP small( &lt; 32*32) and AP medium( &gt;32*32 and &lt;96*96) , <br>\nyolo-v3 got 18.3 and 25.4 scores respective.  <br>\nAnd Centernet got 19.9 and 43 scores respective.<br>\nSo ,if we want  yolo-v3 work, then we should let the feature map bigger than 96*96,  right?",
    "601924": "I agree. Controlling the feature map (grid cell) is one possible solution to detect small object. \n\nThe other way is zooming in by splitting the picture into several parts. In my kernel, I run the detector approximately 4 times for each picture. If you can know the size of the object in advance, this would be the simplest solution.",
    "616388": "qinhui1999 @shivus   I think it could work if first use a region proposal network (like R-CNN) first, or have a single pass network like YOLO produce feature vectors in addition to coordinates that are then passed to the autoencoders. Either way, it could definitely work as one head within a multi-objective network."
  },
  "source": "meta"
}