{
  "id": 126828,
  "title": "In Childhood teachers always told to practice handwriting :(",
  "url": "/competitions/bengaliai-cv19/discussion/126828",
  "author_name": "",
  "post_date": "2020-01-20T14:52:05.173456200Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>A few of the train set example where even I cant understand what is written , its not fair on the model to ask to predict it . \nUnknown Grapheme : \ne.g : What is this Grapheme_Root ? I have no idea . \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F46747f8625f61aef0d5c7013385f8a79%2Fimage.png?generation=1579531673505818&amp;alt=media\" alt=\"\"></p>\n\n<p>Bad Handwriting  : \ne.g. The below one , it can be anything , looks like writer was in a hurry . \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff49a3170370b8bbc4b4f3dce22531827%2Fimage%20(1\" alt=\"\">.png?generation=1579531767150445&amp;alt=media)</p>\n\n<p>Wrong Label in Train Set : \ne.g. The below . Even though model predicted it correctly , it will show validation error .  Question is , do we have wrong labeling in private test set ?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fbb4cd2ce8c0e689d69c53dd820f9238d%2Fimage%20(2\" alt=\"\">.png?generation=1579531894754893&amp;alt=media)</p>",
  "messages": [
    {
      "id": "723866",
      "postDate": "01/20/2020 14:52:05",
      "content": "<p>A few of the train set example where even I cant understand what is written , its not fair on the model to ask to predict it . \nUnknown Grapheme : \ne.g : What is this Grapheme_Root ? I have no idea . \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F46747f8625f61aef0d5c7013385f8a79%2Fimage.png?generation=1579531673505818&amp;alt=media\" alt=\"\"></p>\n\n<p>Bad Handwriting  : \ne.g. The below one , it can be anything , looks like writer was in a hurry . \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff49a3170370b8bbc4b4f3dce22531827%2Fimage%20(1\" alt=\"\">.png?generation=1579531767150445&amp;alt=media)</p>\n\n<p>Wrong Label in Train Set : \ne.g. The below . Even though model predicted it correctly , it will show validation error .  Question is , do we have wrong labeling in private test set ?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fbb4cd2ce8c0e689d69c53dd820f9238d%2Fimage%20(2\" alt=\"\">.png?generation=1579531894754893&amp;alt=media)</p>",
      "rawMarkdown": "A few of the train set example where even I cant understand what is written , its not fair on the model to ask to predict it . \nUnknown Grapheme : \ne.g : What is this Grapheme_Root ? I have no idea . \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F46747f8625f61aef0d5c7013385f8a79%2Fimage.png?generation=1579531673505818&amp;alt=media)\n\nBad Handwriting  : \ne.g. The below one , it can be anything , looks like writer was in a hurry . \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff49a3170370b8bbc4b4f3dce22531827%2Fimage%20(1).png?generation=1579531767150445&amp;alt=media)\n\nWrong Label in Train Set : \ne.g. The below . Even though model predicted it correctly , it will show validation error .  Question is , do we have wrong labeling in private test set ?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fbb4cd2ce8c0e689d69c53dd820f9238d%2Fimage%20(2).png?generation=1579531894754893&amp;alt=media)",
      "votes": null
    },
    {
      "id": "723913",
      "postDate": "01/20/2020 15:45:36",
      "content": "<p>Unfortunately, I think that is the reality of the world as well. Ideally, when deployed data would need to be cleaned to prevent things like this from being (wrongly) classified, but the reality outside of Kaggle is also that there is a lot of noisy data. I think it brings an interesting challenge to also include these noisy elements in the dataset; especially for the hosts, which might get a better understanding of how such a model would perform once deployed in \"the real world\". Just my two cent though.</p>",
      "rawMarkdown": "Unfortunately, I think that is the reality of the world as well. Ideally, when deployed data would need to be cleaned to prevent things like this from being (wrongly) classified, but the reality outside of Kaggle is also that there is a lot of noisy data. I think it brings an interesting challenge to also include these noisy elements in the dataset; especially for the hosts, which might get a better understanding of how such a model would perform once deployed in \"the real world\". Just my two cent though.",
      "votes": null
    },
    {
      "id": "723951",
      "postDate": "01/20/2020 16:25:33",
      "content": "<p>Hey , that's true . Enormous number of people have put effort to collect this data . And my handwriting is nothing better than this :) </p>",
      "rawMarkdown": "Hey , that's true . Enormous number of people have put effort to collect this data . And my handwriting is nothing better than this :)",
      "votes": null
    },
    {
      "id": "723968",
      "postDate": "01/20/2020 16:56:00",
      "content": "<p>Haha, well at least you're able to read this and know when something doesn't look like what it's supposed to! I think a lot of us in this competition don't read / know Bengali and thus just see this as scribbles that don't make any sense. Domain knowledge always helps! ;)</p>",
      "rawMarkdown": "Haha, well at least you're able to read this and know when something doesn't look like what it's supposed to! I think a lot of us in this competition don't read / know Bengali and thus just see this as scribbles that don't make any sense. Domain knowledge always helps! ;)",
      "votes": null
    },
    {
      "id": "724322",
      "postDate": "01/21/2020 03:31:51",
      "content": "<p>I don't think I am being fair to my models at all by asking them to learn what I cant😃 ...the data is pretty overwhelming for a non-native speaker</p>",
      "rawMarkdown": "I don't think I am being fair to my models at all by asking them to learn what I cant😃 ...the data is pretty overwhelming for a non-native speaker",
      "votes": null
    },
    {
      "id": "724323",
      "postDate": "01/21/2020 03:39:25",
      "content": "<p>After the comp is over we can take our models for dinner :P .</p>",
      "rawMarkdown": "After the comp is over we can take our models for dinner :P .",
      "votes": null
    },
    {
      "id": "724370",
      "postDate": "01/21/2020 05:18:34",
      "content": "<p>The first image has not been cropped properly, not handwriting issue. It is looking like 3 but it was cropped from middle.   </p>",
      "rawMarkdown": "The first image has not been cropped properly, not handwriting issue. It is looking like 3 but it was cropped from middle.",
      "votes": null
    },
    {
      "id": "724378",
      "postDate": "01/21/2020 05:30:33",
      "content": "<p>humm .. that is a good observation .. Thanks Sourin.</p>",
      "rawMarkdown": "humm .. that is a good observation .. Thanks Sourin.",
      "votes": null
    },
    {
      "id": "724445",
      "postDate": "01/21/2020 06:49:28",
      "content": "<p>Is there is any method to remove these datasets?? I think it will help us increase the accuracy</p>",
      "rawMarkdown": "Is there is any method to remove these datasets?? I think it will help us increase the accuracy",
      "votes": null
    },
    {
      "id": "724451",
      "postDate": "01/21/2020 06:51:57",
      "content": "<p>Good job &amp;resolution</p>",
      "rawMarkdown": "Good job &amp;resolution",
      "votes": null
    },
    {
      "id": "724538",
      "postDate": "01/21/2020 08:20:36",
      "content": "<p>Maybe on training accuracy, but not necessarily on the public / private leaderboard. Unless you find a way to automatically remove these, but then you've made a classifier, which is exactly what we're trying to do in the first place ;)</p>\n\n<p>I think having some noise is also a way to prevent overfitting and can in some measure maybe be used as regularization. Also, you want your training dataset to be as close as possible to your testing one, and we can be fairly confident there are going to be noisy images in the leaderboard datasets.</p>",
      "rawMarkdown": "Maybe on training accuracy, but not necessarily on the public / private leaderboard. Unless you find a way to automatically remove these, but then you've made a classifier, which is exactly what we're trying to do in the first place ;)\n\nI think having some noise is also a way to prevent overfitting and can in some measure maybe be used as regularization. Also, you want your training dataset to be as close as possible to your testing one, and we can be fairly confident there are going to be noisy images in the leaderboard datasets.",
      "votes": null
    },
    {
      "id": "724556",
      "postDate": "01/21/2020 08:40:08",
      "content": "<p>I think thats correct suggestion <a href=\"/maxlenormand\">@maxlenormand</a> . It works as augmentation in training .</p>",
      "rawMarkdown": "I think thats correct suggestion @maxlenormand . It works as augmentation in training .",
      "votes": null
    },
    {
      "id": "724599",
      "postDate": "01/21/2020 09:31:38",
      "content": "<p>Some images even humans cannot classify.\nEven though I know the language it was difficult for me to classify just by seeing some of them (Example above picture 2). So it is obvious that machines will make mistakes while classifying. But if more data is there on similar lines I think the accuracy will improve automatically.</p>",
      "rawMarkdown": "Some images even humans cannot classify.\nEven though I know the language it was difficult for me to classify just by seeing some of them (Example above picture 2). So it is obvious that machines will make mistakes while classifying. But if more data is there on similar lines I think the accuracy will improve automatically.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 723913,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "01/20/2020 15:45:36",
      "content": "<p>Unfortunately, I think that is the reality of the world as well. Ideally, when deployed data would need to be cleaned to prevent things like this from being (wrongly) classified, but the reality outside of Kaggle is also that there is a lot of noisy data. I think it brings an interesting challenge to also include these noisy elements in the dataset; especially for the hosts, which might get a better understanding of how such a model would perform once deployed in \"the real world\". Just my two cent though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 723951,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/20/2020 16:25:33",
          "content": "<p>Hey , that's true . Enormous number of people have put effort to collect this data . And my handwriting is nothing better than this :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 723968,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "01/20/2020 16:56:00",
          "content": "<p>Haha, well at least you're able to read this and know when something doesn't look like what it's supposed to! I think a lot of us in this competition don't read / know Bengali and thus just see this as scribbles that don't make any sense. Domain knowledge always helps! ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 724322,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "01/21/2020 03:31:51",
      "content": "<p>I don't think I am being fair to my models at all by asking them to learn what I cant😃 ...the data is pretty overwhelming for a non-native speaker</p>",
      "votes": null,
      "replies": [
        {
          "id": 724323,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/21/2020 03:39:25",
          "content": "<p>After the comp is over we can take our models for dinner :P .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 724370,
      "author_name": "sourinkarmakar",
      "author_url": "",
      "post_date": "01/21/2020 05:18:34",
      "content": "<p>The first image has not been cropped properly, not handwriting issue. It is looking like 3 but it was cropped from middle.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 724378,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/21/2020 05:30:33",
          "content": "<p>humm .. that is a good observation .. Thanks Sourin.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 724445,
      "author_name": "chekoduadarsh",
      "author_url": "",
      "post_date": "01/21/2020 06:49:28",
      "content": "<p>Is there is any method to remove these datasets?? I think it will help us increase the accuracy</p>",
      "votes": null,
      "replies": [
        {
          "id": 724538,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "01/21/2020 08:20:36",
          "content": "<p>Maybe on training accuracy, but not necessarily on the public / private leaderboard. Unless you find a way to automatically remove these, but then you've made a classifier, which is exactly what we're trying to do in the first place ;)</p>\n\n<p>I think having some noise is also a way to prevent overfitting and can in some measure maybe be used as regularization. Also, you want your training dataset to be as close as possible to your testing one, and we can be fairly confident there are going to be noisy images in the leaderboard datasets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724556,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/21/2020 08:40:08",
          "content": "<p>I think thats correct suggestion <a href=\"/maxlenormand\">@maxlenormand</a> . It works as augmentation in training .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 724599,
          "author_name": "sourinkarmakar",
          "author_url": "",
          "post_date": "01/21/2020 09:31:38",
          "content": "<p>Some images even humans cannot classify.\nEven though I know the language it was difficult for me to classify just by seeing some of them (Example above picture 2). So it is obvious that machines will make mistakes while classifying. But if more data is there on similar lines I think the accuracy will improve automatically.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 724451,
      "author_name": "",
      "author_url": "",
      "post_date": "01/21/2020 06:51:57",
      "content": "<p>Good job &amp;resolution</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "723866": "A few of the train set example where even I cant understand what is written , its not fair on the model to ask to predict it . \nUnknown Grapheme : \ne.g : What is this Grapheme_Root ? I have no idea . \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F46747f8625f61aef0d5c7013385f8a79%2Fimage.png?generation=1579531673505818&amp;alt=media)\n\nBad Handwriting  : \ne.g. The below one , it can be anything , looks like writer was in a hurry . \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Ff49a3170370b8bbc4b4f3dce22531827%2Fimage%20(1).png?generation=1579531767150445&amp;alt=media)\n\nWrong Label in Train Set : \ne.g. The below . Even though model predicted it correctly , it will show validation error .  Question is , do we have wrong labeling in private test set ?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fbb4cd2ce8c0e689d69c53dd820f9238d%2Fimage%20(2).png?generation=1579531894754893&amp;alt=media)",
    "723913": "Unfortunately, I think that is the reality of the world as well. Ideally, when deployed data would need to be cleaned to prevent things like this from being (wrongly) classified, but the reality outside of Kaggle is also that there is a lot of noisy data. I think it brings an interesting challenge to also include these noisy elements in the dataset; especially for the hosts, which might get a better understanding of how such a model would perform once deployed in \"the real world\". Just my two cent though.",
    "723951": "Hey , that's true . Enormous number of people have put effort to collect this data . And my handwriting is nothing better than this :)",
    "723968": "Haha, well at least you're able to read this and know when something doesn't look like what it's supposed to! I think a lot of us in this competition don't read / know Bengali and thus just see this as scribbles that don't make any sense. Domain knowledge always helps! ;)",
    "724322": "I don't think I am being fair to my models at all by asking them to learn what I cant😃 ...the data is pretty overwhelming for a non-native speaker",
    "724323": "After the comp is over we can take our models for dinner :P .",
    "724370": "The first image has not been cropped properly, not handwriting issue. It is looking like 3 but it was cropped from middle.",
    "724378": "humm .. that is a good observation .. Thanks Sourin.",
    "724445": "Is there is any method to remove these datasets?? I think it will help us increase the accuracy",
    "724451": "Good job &amp;resolution",
    "724538": "Maybe on training accuracy, but not necessarily on the public / private leaderboard. Unless you find a way to automatically remove these, but then you've made a classifier, which is exactly what we're trying to do in the first place ;)\n\nI think having some noise is also a way to prevent overfitting and can in some measure maybe be used as regularization. Also, you want your training dataset to be as close as possible to your testing one, and we can be fairly confident there are going to be noisy images in the leaderboard datasets.",
    "724556": "I think thats correct suggestion @maxlenormand . It works as augmentation in training .",
    "724599": "Some images even humans cannot classify.\nEven though I know the language it was difficult for me to classify just by seeing some of them (Example above picture 2). So it is obvious that machines will make mistakes while classifying. But if more data is there on similar lines I think the accuracy will improve automatically."
  },
  "source": "meta"
}