{
  "id": 126202,
  "title": "report alternative approach here",
  "url": "/competitions/bengaliai-cv19/discussion/126202",
  "author_name": "hengck23",
  "post_date": "2020-01-16T08:58:42.251000",
  "votes": 7,
  "comment_count": 16,
  "views": 0,
  "content": "<p>the usual approach is treat the problem as image classification.\nif you are other approaches, you can share the results here.</p>\n\n<p>examples include:</p>\n\n<ol>\n<li><p>object instance detection (i.e. you have a bounding box annotation for the text (or constant, vowel part) and take approach like faster-rcnn)</p></li>\n<li><p>object instance segmentation (i.e. you have a mask annotation for the text (or part) and take approach like mask-rcnn)</p></li>\n<li><p>semantic segmentation (i.e. you pixel label the text (or part) and take approach like unet)</p></li>\n</ol>",
  "messages": [
    {
      "id": 720239,
      "postDate": "2020-01-16T08:58:42.250Z",
      "content": "<p>the usual approach is treat the problem as image classification.\nif you are other approaches, you can share the results here.</p>\n\n<p>examples include:</p>\n\n<ol>\n<li><p>object instance detection (i.e. you have a bounding box annotation for the text (or constant, vowel part) and take approach like faster-rcnn)</p></li>\n<li><p>object instance segmentation (i.e. you have a mask annotation for the text (or part) and take approach like mask-rcnn)</p></li>\n<li><p>semantic segmentation (i.e. you pixel label the text (or part) and take approach like unet)</p></li>\n</ol>",
      "rawMarkdown": "the usual approach is treat the problem as image classification.\nif you are other approaches, you can share the results here.\n\nexamples include:\n\n1. object instance detection (i.e. you have a bounding box annotation for the text (or constant, vowel part) and take approach like faster-rcnn)\n\n2. object instance segmentation (i.e. you have a mask annotation for the text (or part) and take approach like mask-rcnn)\n\n3. semantic segmentation (i.e. you pixel label the text (or part) and take approach like unet)\n\n",
      "votes": 7
    },
    {
      "id": 720734,
      "postDate": "2020-01-16T17:23:17.500Z",
      "content": "<p>Hi Heng , \nI have not started it yet , because it seemed complex . But I am planning to try\n Siamese network like Humpback Whale . I have no idea about it for now as to how to find anchors etc .But seemed to be something new to try . </p>\n\n<p>There is another thing I want to try is the Ekush dataset addition with our current data for rare classes . Only work is , the classes are different between both dataset so I need to see the letters and reannotate them before feeding . Was thinking to use them as pseudolabel as well .</p>",
      "rawMarkdown": "Hi Heng , \nI have not started it yet , because it seemed complex . But I am planning to try\n Siamese network like Humpback Whale . I have no idea about it for now as to how to find anchors etc .But seemed to be something new to try . \n\nThere is another thing I want to try is the Ekush dataset addition with our current data for rare classes . Only work is , the classes are different between both dataset so I need to see the letters and reannotate them before feeding . Was thinking to use them as pseudolabel as well .\n",
      "votes": 1
    },
    {
      "id": 720311,
      "postDate": "2020-01-16T10:11:21.753Z",
      "content": "<p>Hmm, but wouldn't you need masks to do those segmentation tasks?</p>\n\n<p>Making classification from segmentation tasks is a very interesting idea (<a href=\"/cdeotte\">@cdeotte</a> did something along those lines during the <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\">Understanding clouds competitions</a>) but the other way around? I'm curious.</p>\n\n<p>I guess that would require a bit more domain knowledge, use OpenCV to keep the written parts and then separate them with a bit of manual work? Or otherwise use heatmaps of the areas the model is using to make it's prediction?</p>\n\n<p>Either way, I'd be curious to see how it would be done!</p>",
      "rawMarkdown": "Hmm, but wouldn't you need masks to do those segmentation tasks?\n\nMaking classification from segmentation tasks is a very interesting idea (@cdeotte did something along those lines during the [Understanding clouds competitions](https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60)) but the other way around? I'm curious.\n\nI guess that would require a bit more domain knowledge, use OpenCV to keep the written parts and then separate them with a bit of manual work? Or otherwise use heatmaps of the areas the model is using to make it's prediction?\n\nEither way, I'd be curious to see how it would be done!",
      "votes": 1,
      "replies": [
        {
          "id": 720740,
          "postDate": "2020-01-16T17:30:24.743Z",
          "content": "<p>Check this out for bounding box model where I believe masks were not given \n<a href=\"https://www.kaggle.com/martinpiotte/bounding-box-model\">https://www.kaggle.com/martinpiotte/bounding-box-model</a></p>",
          "rawMarkdown": "Check this out for bounding box model where I believe masks were not given \nhttps://www.kaggle.com/martinpiotte/bounding-box-model"
        },
        {
          "id": 720899,
          "postDate": "2020-01-16T20:46:16.057Z",
          "content": "<p>Interesting, I already did something <a href=\"https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images\">a bit similar</a> but that doesn't resolve the problem. Here our images are composed of 3 labels, and what I was saying is that we would (probably) want to make 3 different masks for the root, vowel and consonant in each character.</p>\n\n<p>My question was more about, how can we get that info? I barely know anything about these characters; domain knowledge would definitely help for this.</p>\n\n<p>That's what I was asking. If you have any ideas, please feel free to share, I'd be curious to hear them!</p>",
          "rawMarkdown": "Interesting, I already did something [a bit similar](https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images) but that doesn't resolve the problem. Here our images are composed of 3 labels, and what I was saying is that we would (probably) want to make 3 different masks for the root, vowel and consonant in each character.\n\nMy question was more about, how can we get that info? I barely know anything about these characters; domain knowledge would definitely help for this.\n\nThat's what I was asking. If you have any ideas, please feel free to share, I'd be curious to hear them!",
          "votes": 1
        },
        {
          "id": 721237,
          "postDate": "2020-01-17T07:51:51.590Z",
          "content": "<p>Has anyone tried to look at attention maps (using GradCam or anything else) for the three different heads of a basic/good model to see if each head is focusing on a different part of the image?\nTo see if a trained model without any domain knowledge could actually give domain knowledge? (I mean I have no idea which part is which when looking at an image but the model probably learnt it somehow so it could teach that for free)</p>",
          "rawMarkdown": "Has anyone tried to look at attention maps (using GradCam or anything else) for the three different heads of a basic/good model to see if each head is focusing on a different part of the image?\nTo see if a trained model without any domain knowledge could actually give domain knowledge? (I mean I have no idea which part is which when looking at an image but the model probably learnt it somehow so it could teach that for free)"
        },
        {
          "id": 721243,
          "postDate": "2020-01-17T08:06:20.447Z",
          "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a>  . Okay I got it now . I have the domain  knowledge as my mother tongue is Bengali . But do you know any effective method to create mask around the letters automatically ? There are huge number of characters . How , would you approach creating bounding boxes around them ?</p>",
          "rawMarkdown": "@maxlenormand  . Okay I got it now . I have the domain  knowledge as my mother tongue is Bengali . But do you know any effective method to create mask around the letters automatically ? There are huge number of characters . How , would you approach creating bounding boxes around them ?"
        },
        {
          "id": 721326,
          "postDate": "2020-01-17T09:24:26.180Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> this is what I was thinking about in my first comment :)\nI haven't ever implemented it, but that would definitely be an interesting thing to look into, even more for error analysis and trying to improve predictions!</p>\n\n<p><a href=\"/phoenix9032\">@phoenix9032</a> That's exactly the problem, I don't quite know how to do this at scale. Maybe one approach could be to manually mask some of the three different components at first, then perform unsupervised classifications to see if a model could separate the 3 characters from each other (though I doubt that it would produce usable results). If we can separate them easily, we can then create masks using OpenCV or similar, by simply threshold the pixels above a certain value.</p>\n\n<p>Does that make sense?</p>\n\n<p>Either way, it would be quite a hustle to implement, but I've done something similar for work in the past.</p>",
          "rawMarkdown": "@optimo this is what I was thinking about in my first comment :)\nI haven't ever implemented it, but that would definitely be an interesting thing to look into, even more for error analysis and trying to improve predictions!\n\n@phoenix9032 That's exactly the problem, I don't quite know how to do this at scale. Maybe one approach could be to manually mask some of the three different components at first, then perform unsupervised classifications to see if a model could separate the 3 characters from each other (though I doubt that it would produce usable results). If we can separate them easily, we can then create masks using OpenCV or similar, by simply threshold the pixels above a certain value.\n\nDoes that make sense?\n\nEither way, it would be quite a hustle to implement, but I've done something similar for work in the past."
        },
        {
          "id": 721366,
          "postDate": "2020-01-17T10:04:07.810Z",
          "content": "<p>you probably don't have to do for all since current approach already get &gt;96% accurate. just hand label those that are wrong</p>",
          "rawMarkdown": "you probably don't have to do for all since current approach already get &gt;96% accurate. just hand label those that are wrong",
          "votes": 1
        },
        {
          "id": 721380,
          "postDate": "2020-01-17T10:26:58.823Z",
          "content": "<p>Okay got it ..</p>",
          "rawMarkdown": "Okay got it .."
        },
        {
          "id": 721417,
          "postDate": "2020-01-17T11:02:58.447Z",
          "content": "<p>Fair point.\nThough there are also other combinations in the private test set, maybe one would need to do a bit more?</p>\n\n<p>I don't think I'd have the time to implement that, but I would be really curious in seeing what results it would produce!</p>",
          "rawMarkdown": "Fair point.\nThough there are also other combinations in the private test set, maybe one would need to do a bit more?\n\nI don't think I'd have the time to implement that, but I would be really curious in seeing what results it would produce!"
        },
        {
          "id": 721676,
          "postDate": "2020-01-17T15:44:08.833Z",
          "content": "<p>One more hint: just hand label the weakest class. Or create more data, eg ask your friends to contribute their hand writing </p>",
          "rawMarkdown": "One more hint: just hand label the weakest class. Or create more data, eg ask your friends to contribute their hand writing ",
          "votes": 1
        },
        {
          "id": 721742,
          "postDate": "2020-01-17T16:53:15.463Z",
          "content": "<p>Your \"friends\" better not be competing separately though . That would be private sharing .  Just a word of caution.</p>",
          "rawMarkdown": "Your \"friends\" better not be competing separately though . That would be private sharing .  Just a word of caution.",
          "votes": 1
        },
        {
          "id": 721825,
          "postDate": "2020-01-17T18:29:46.773Z",
          "content": "<p>Per the rules, any extra hand writing samples would need to be made available to all competitors and disclosed in the external data thread. With that said, we'd be thrilled to see more data contributed!</p>",
          "rawMarkdown": "Per the rules, any extra hand writing samples would need to be made available to all competitors and disclosed in the external data thread. With that said, we'd be thrilled to see more data contributed!",
          "votes": 4
        },
        {
          "id": 722067,
          "postDate": "2020-01-18T03:53:38.940Z",
          "content": "<p>Hi Heng,\nWhat would be a good way to create masks?</p>",
          "rawMarkdown": "Hi Heng,\nWhat would be a good way to create masks?\n"
        },
        {
          "id": 722278,
          "postDate": "2020-01-18T10:53:54.723Z",
          "content": "<p>It will be very nice , if some competitor can fill the A4 sheet in the data repo with their handwriting , It might be interesting to see how a person with different ethnicity with recently learnt language perceive and  write Bengali letters .</p>",
          "rawMarkdown": "It will be very nice , if some competitor can fill the A4 sheet in the data repo with their handwriting , It might be interesting to see how a person with different ethnicity with recently learnt language perceive and  write Bengali letters ."
        }
      ]
    },
    {
      "id": 721919,
      "postDate": "2020-01-17T21:49:51.107Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 720734,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2020-01-16T17:23:17.500000",
      "content": "<p>Hi Heng , \nI have not started it yet , because it seemed complex . But I am planning to try\n Siamese network like Humpback Whale . I have no idea about it for now as to how to find anchors etc .But seemed to be something new to try . </p>\n\n<p>There is another thing I want to try is the Ekush dataset addition with our current data for rare classes . Only work is , the classes are different between both dataset so I need to see the letters and reannotate them before feeding . Was thinking to use them as pseudolabel as well .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 720311,
      "author_name": "Maxime Lenormand",
      "author_url": "",
      "post_date": "2020-01-16T10:11:21.753000",
      "content": "<p>Hmm, but wouldn't you need masks to do those segmentation tasks?</p>\n\n<p>Making classification from segmentation tasks is a very interesting idea (<a href=\"/cdeotte\">@cdeotte</a> did something along those lines during the <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\">Understanding clouds competitions</a>) but the other way around? I'm curious.</p>\n\n<p>I guess that would require a bit more domain knowledge, use OpenCV to keep the written parts and then separate them with a bit of manual work? Or otherwise use heatmaps of the areas the model is using to make it's prediction?</p>\n\n<p>Either way, I'd be curious to see how it would be done!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 720740,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-16T17:30:24.743000",
          "content": "<p>Check this out for bounding box model where I believe masks were not given \n<a href=\"https://www.kaggle.com/martinpiotte/bounding-box-model\">https://www.kaggle.com/martinpiotte/bounding-box-model</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720899,
          "author_name": "Maxime Lenormand",
          "author_url": "",
          "post_date": "2020-01-16T20:46:16.057000",
          "content": "<p>Interesting, I already did something <a href=\"https://www.kaggle.com/maxlenormand/cropping-to-character-resizing-images\">a bit similar</a> but that doesn't resolve the problem. Here our images are composed of 3 labels, and what I was saying is that we would (probably) want to make 3 different masks for the root, vowel and consonant in each character.</p>\n\n<p>My question was more about, how can we get that info? I barely know anything about these characters; domain knowledge would definitely help for this.</p>\n\n<p>That's what I was asking. If you have any ideas, please feel free to share, I'd be curious to hear them!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721237,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-01-17T07:51:51.590000",
          "content": "<p>Has anyone tried to look at attention maps (using GradCam or anything else) for the three different heads of a basic/good model to see if each head is focusing on a different part of the image?\nTo see if a trained model without any domain knowledge could actually give domain knowledge? (I mean I have no idea which part is which when looking at an image but the model probably learnt it somehow so it could teach that for free)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721243,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-17T08:06:20.447000",
          "content": "<p><a href=\"/maxlenormand\">@maxlenormand</a>  . Okay I got it now . I have the domain  knowledge as my mother tongue is Bengali . But do you know any effective method to create mask around the letters automatically ? There are huge number of characters . How , would you approach creating bounding boxes around them ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721326,
          "author_name": "Maxime Lenormand",
          "author_url": "",
          "post_date": "2020-01-17T09:24:26.180000",
          "content": "<p><a href=\"/optimo\">@optimo</a> this is what I was thinking about in my first comment :)\nI haven't ever implemented it, but that would definitely be an interesting thing to look into, even more for error analysis and trying to improve predictions!</p>\n\n<p><a href=\"/phoenix9032\">@phoenix9032</a> That's exactly the problem, I don't quite know how to do this at scale. Maybe one approach could be to manually mask some of the three different components at first, then perform unsupervised classifications to see if a model could separate the 3 characters from each other (though I doubt that it would produce usable results). If we can separate them easily, we can then create masks using OpenCV or similar, by simply threshold the pixels above a certain value.</p>\n\n<p>Does that make sense?</p>\n\n<p>Either way, it would be quite a hustle to implement, but I've done something similar for work in the past.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721366,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-17T10:04:07.810000",
          "content": "<p>you probably don't have to do for all since current approach already get &gt;96% accurate. just hand label those that are wrong</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721380,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-17T10:26:58.823000",
          "content": "<p>Okay got it ..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721417,
          "author_name": "Maxime Lenormand",
          "author_url": "",
          "post_date": "2020-01-17T11:02:58.447000",
          "content": "<p>Fair point.\nThough there are also other combinations in the private test set, maybe one would need to do a bit more?</p>\n\n<p>I don't think I'd have the time to implement that, but I would be really curious in seeing what results it would produce!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 721676,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-17T15:44:08.833000",
          "content": "<p>One more hint: just hand label the weakest class. Or create more data, eg ask your friends to contribute their hand writing </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721742,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-17T16:53:15.463000",
          "content": "<p>Your \"friends\" better not be competing separately though . That would be private sharing .  Just a word of caution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721825,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-01-17T18:29:46.773000",
          "content": "<p>Per the rules, any extra hand writing samples would need to be made available to all competitors and disclosed in the external data thread. With that said, we'd be thrilled to see more data contributed!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 722067,
          "author_name": "Akhilesh",
          "author_url": "",
          "post_date": "2020-01-18T03:53:38.940000",
          "content": "<p>Hi Heng,\nWhat would be a good way to create masks?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722278,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-18T10:53:54.723000",
          "content": "<p>It will be very nice , if some competitor can fill the A4 sheet in the data repo with their handwriting , It might be interesting to see how a person with different ethnicity with recently learnt language perceive and  write Bengali letters .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 721919,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-17T21:49:51.107000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "720239": "the usual approach is treat the problem as image classification.\nif you are other approaches, you can share the results here.\n\nexamples include:\n\n1. object instance detection (i.e. you have a bounding box annotation for the text (or constant, vowel part) and take approach like faster-rcnn)\n\n2. object instance segmentation (i.e. you have a mask annotation for the text (or part) and take approach like mask-rcnn)\n\n3. semantic segmentation (i.e. you pixel label the text (or part) and take approach like unet)\n\n",
    "720734": "Hi Heng , \nI have not started it yet , because it seemed complex . But I am planning to try\n Siamese network like Humpback Whale . I have no idea about it for now as to how to find anchors etc .But seemed to be something new to try . \n\nThere is another thing I want to try is the Ekush dataset addition with our current data for rare classes . Only work is , the classes are different between both dataset so I need to see the letters and reannotate them before feeding . Was thinking to use them as pseudolabel as well .\n",
    "720311": "Hmm, but wouldn't you need masks to do those segmentation tasks?\n\nMaking classification from segmentation tasks is a very interesting idea (@cdeotte did something along those lines during the [Understanding clouds competitions](https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60)) but the other way around? I'm curious.\n\nI guess that would require a bit more domain knowledge, use OpenCV to keep the written parts and then separate them with a bit of manual work? Or otherwise use heatmaps of the areas the model is using to make it's prediction?\n\nEither way, I'd be curious to see how it would be done!",
    "721919": ""
  }
}