{
  "id": 127763,
  "title": "[Domain knowledge needed] Error analysis",
  "url": "/competitions/bengaliai-cv19/discussion/127763",
  "author_name": "",
  "post_date": "2020-01-26T14:24:19.618003300Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi! Been striving to push single model past .971 without using computationally-expensive stuff. When analyzing the confusion matrix I found an interesting pattern. Here is the confusion matrix for my model (validation, grapheme root).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F85dcbbbb4b353632296b09b8155ca419%2FConfusion.png?generation=1580048544322053&amp;alt=media\" alt=\"\"></p>\n\n<p>I have zeroed out the diagonal and normalized according to frequency for each class.\nIt is quite evident that there is a bright band very close to the diagonal. Does that mean grapheme roots with near IDs are very similar?</p>",
  "messages": [
    {
      "id": "729687",
      "postDate": "01/26/2020 14:24:19",
      "content": "<p>Hi! Been striving to push single model past .971 without using computationally-expensive stuff. When analyzing the confusion matrix I found an interesting pattern. Here is the confusion matrix for my model (validation, grapheme root).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F85dcbbbb4b353632296b09b8155ca419%2FConfusion.png?generation=1580048544322053&amp;alt=media\" alt=\"\"></p>\n\n<p>I have zeroed out the diagonal and normalized according to frequency for each class.\nIt is quite evident that there is a bright band very close to the diagonal. Does that mean grapheme roots with near IDs are very similar?</p>",
      "rawMarkdown": "Hi! Been striving to push single model past .971 without using computationally-expensive stuff. When analyzing the confusion matrix I found an interesting pattern. Here is the confusion matrix for my model (validation, grapheme root).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F85dcbbbb4b353632296b09b8155ca419%2FConfusion.png?generation=1580048544322053&amp;alt=media)\n\nI have zeroed out the diagonal and normalized according to frequency for each class.\nIt is quite evident that there is a bright band very close to the diagonal. Does that mean grapheme roots with near IDs are very similar?",
      "votes": null
    },
    {
      "id": "729865",
      "postDate": "01/26/2020 18:22:22",
      "content": "<p>I didn't understand the confusion matrix properly . But to answer your question . Yes, grapheme roots are mostly ordered based on their base grapheme root . Like sound \"K\" is the base grapheme_root . Now you can make compound alphabets (grapheme_roots) using other base alphabets with it . Ex. KK , KB , KT , KM . Now the IDs of K,KK,KB,KT,KM all will be in order and since the base grapheme root is same , it will look similar too .</p>",
      "rawMarkdown": "I didn't understand the confusion matrix properly . But to answer your question . Yes, grapheme roots are mostly ordered based on their base grapheme root . Like sound \"K\" is the base grapheme_root . Now you can make compound alphabets (grapheme_roots) using other base alphabets with it . Ex. KK , KB , KT , KM . Now the IDs of K,KK,KB,KT,KM all will be in order and since the base grapheme root is same , it will look similar too .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 729865,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "01/26/2020 18:22:22",
      "content": "<p>I didn't understand the confusion matrix properly . But to answer your question . Yes, grapheme roots are mostly ordered based on their base grapheme root . Like sound \"K\" is the base grapheme_root . Now you can make compound alphabets (grapheme_roots) using other base alphabets with it . Ex. KK , KB , KT , KM . Now the IDs of K,KK,KB,KT,KM all will be in order and since the base grapheme root is same , it will look similar too .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "729687": "Hi! Been striving to push single model past .971 without using computationally-expensive stuff. When analyzing the confusion matrix I found an interesting pattern. Here is the confusion matrix for my model (validation, grapheme root).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1943421%2F85dcbbbb4b353632296b09b8155ca419%2FConfusion.png?generation=1580048544322053&amp;alt=media)\n\nI have zeroed out the diagonal and normalized according to frequency for each class.\nIt is quite evident that there is a bright band very close to the diagonal. Does that mean grapheme roots with near IDs are very similar?",
    "729865": "I didn't understand the confusion matrix properly . But to answer your question . Yes, grapheme roots are mostly ordered based on their base grapheme root . Like sound \"K\" is the base grapheme_root . Now you can make compound alphabets (grapheme_roots) using other base alphabets with it . Ex. KK , KB , KT , KM . Now the IDs of K,KK,KB,KT,KM all will be in order and since the base grapheme root is same , it will look similar too ."
  },
  "source": "meta"
}