{
  "id": 133568,
  "title": "What do you think of the shake up?",
  "url": "/competitions/bengaliai-cv19/discussion/133568",
  "author_name": "",
  "post_date": "2020-03-03T10:26:43.631770100Z",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Two of my submissions:\n1) local 0.9968, LB 0.9884\n2) local 0.9972, LB 0.9881\nthe local - PB gap for my model (same model, same fold) ranges from 0.0083 to 0.0091\nthus I guess there may be ~0.0007 shake up for a single model (single fold)</p>\n\n<p>---Update---\nlocal 0.9978, LB 0.9895, my gap comes back into 0.0083</p>",
  "messages": [
    {
      "id": "762213",
      "postDate": "03/03/2020 10:26:43",
      "content": "<p>Two of my submissions:\n1) local 0.9968, LB 0.9884\n2) local 0.9972, LB 0.9881\nthe local - PB gap for my model (same model, same fold) ranges from 0.0083 to 0.0091\nthus I guess there may be ~0.0007 shake up for a single model (single fold)</p>\n\n<p>---Update---\nlocal 0.9978, LB 0.9895, my gap comes back into 0.0083</p>",
      "rawMarkdown": "Two of my submissions:\n1) local 0.9968, LB 0.9884\n2) local 0.9972, LB 0.9881\nthe local - PB gap for my model (same model, same fold) ranges from 0.0083 to 0.0091\nthus I guess there may be ~0.0007 shake up for a single model (single fold)\n\n---Update---\nlocal 0.9978, LB 0.9895, my gap comes back into 0.0083",
      "votes": null
    },
    {
      "id": "762226",
      "postDate": "03/03/2020 10:42:20",
      "content": "<p>glad you brought this topic. I'm also experiencing something similar. For the most part, CV and LB correlates well but there are times when local CV doesn't correlate well with LB; don't know why?\n```\nCV: 996409\nLB: 0.9889</p>\n\n<p>CV: 996608\nLB: 0.9883\n```\nEDIT: So yes, I also expect little shakeup :(</p>",
      "rawMarkdown": "glad you brought this topic. I'm also experiencing something similar. For the most part, CV and LB correlates well but there are times when local CV doesn't correlate well with LB; don't know why?\n```\nCV: 996409\nLB: 0.9889\n\nCV: 996608\nLB: 0.9883\n```\nEDIT: So yes, I also expect little shakeup :(",
      "votes": null
    },
    {
      "id": "762238",
      "postDate": "03/03/2020 11:02:11",
      "content": "<p>Not huge I guess. If only there are really unuasual privat set words, like from old papers and libraries</p>",
      "rawMarkdown": "Not huge I guess. If only there are really unuasual privat set words, like from old papers and libraries",
      "votes": null
    },
    {
      "id": "762373",
      "postDate": "03/03/2020 13:33:32",
      "content": "<p>My concern is where unseen graphemes are. If unseen graphemes are only in private test set, there would be a shake.\nI'm not sure how many unseen graphemes actually exist. The possible combinations are 12936 (168 * 11 * 7). There are 1292 graphemes in training set.\nThe <a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">slide</a> says that they selected 1295 commonly used bengali graphemes...</p>",
      "rawMarkdown": "My concern is where unseen graphemes are. If unseen graphemes are only in private test set, there would be a shake.\nI'm not sure how many unseen graphemes actually exist. The possible combinations are 12936 (168 * 11 * 7). There are 1292 graphemes in training set.\nThe [slide](https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf) says that they selected 1295 commonly used bengali graphemes...",
      "votes": null
    },
    {
      "id": "762439",
      "postDate": "03/03/2020 14:08:35",
      "content": "<p>I guess all the unseen graphemes in the private LB. If they are split to both public and private, it is hard to achieve 0.99+. </p>",
      "rawMarkdown": "I guess all the unseen graphemes in the private LB. If they are split to both public and private, it is hard to achieve 0.99+.",
      "votes": null
    },
    {
      "id": "762620",
      "postDate": "03/03/2020 16:48:45",
      "content": "<p>no worry. do this test:</p>\n\n<ol>\n<li>train a classifier using 1295 grapheme class</li>\n<li>use this to decode into root, vowel, constant</li>\n<li>make a submission</li>\n</ol>\n\n<p>the cv/lb gap tells you  how many \"unseen graphemes\" there are ...\n(more correctly it tells you how many \"unidentifiable graphemes\"  there are )</p>\n\n<p>it think it is not a lot </p>",
      "rawMarkdown": "no worry. do this test:\n\n1. train a classifier using 1295 grapheme class\n2. use this to decode into root, vowel, constant\n3. make a submission\n\nthe cv/lb gap tells you  how many \"unseen graphemes\" there are ...\n(more correctly it tells you how many \"unidentifiable graphemes\"  there are )\n\n it think it is not a lot",
      "votes": null
    },
    {
      "id": "762961",
      "postDate": "03/04/2020 01:23:53",
      "content": "<p>Same for me, do observe some differences. I think this might be particularly alarming for such a dense LB. There can be 15 people in that range....\nLocal .9965: LB .9887\nLocal .9967; LB .9885</p>\n\n<p>Are those two experiments from the same model in your cases? My experiments above are from two (although very similar) models</p>",
      "rawMarkdown": "Same for me, do observe some differences. I think this might be particularly alarming for such a dense LB. There can be 15 people in that range....\nLocal .9965: LB .9887\nLocal .9967; LB .9885\n\nAre those two experiments from the same model in your cases? My experiments above are from two (although very similar) models",
      "votes": null
    },
    {
      "id": "762985",
      "postDate": "03/04/2020 01:54:45",
      "content": "<p>You have too much gap between CV and LB😂 </p>",
      "rawMarkdown": "You have too much gap between CV and LB😂",
      "votes": null
    },
    {
      "id": "762991",
      "postDate": "03/04/2020 02:07:35",
      "content": "<p>yes same model ~</p>",
      "rawMarkdown": "yes same model ~",
      "votes": null
    },
    {
      "id": "763000",
      "postDate": "03/04/2020 02:28:53",
      "content": "<p><a href=\"/murphy89\">@murphy89</a> I actually thought the same. But it seems the gap becomes greater the more your CV gets to 1.00😧 Had a small CV/LB gap myself of consistent ~.007</p>",
      "rawMarkdown": "murphy89 I actually thought the same. But it seems the gap becomes greater the more your CV gets to 1.00😧 Had a small CV/LB gap myself of consistent ~.007",
      "votes": null
    },
    {
      "id": "764335",
      "postDate": "03/05/2020 11:05:05",
      "content": "<p>How do you think,  should we trust cv or lb more? (Which of these models you`d choose as final score on private)</p>",
      "rawMarkdown": "How do you think,  should we trust cv or lb more? (Which of these models you`d choose as final score on private)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 762226,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "03/03/2020 10:42:20",
      "content": "<p>glad you brought this topic. I'm also experiencing something similar. For the most part, CV and LB correlates well but there are times when local CV doesn't correlate well with LB; don't know why?\n```\nCV: 996409\nLB: 0.9889</p>\n\n<p>CV: 996608\nLB: 0.9883\n```\nEDIT: So yes, I also expect little shakeup :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 762985,
          "author_name": "murphy89",
          "author_url": "",
          "post_date": "03/04/2020 01:54:45",
          "content": "<p>You have too much gap between CV and LB😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763000,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/04/2020 02:28:53",
          "content": "<p><a href=\"/murphy89\">@murphy89</a> I actually thought the same. But it seems the gap becomes greater the more your CV gets to 1.00😧 Had a small CV/LB gap myself of consistent ~.007</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764335,
          "author_name": "kupchanski",
          "author_url": "",
          "post_date": "03/05/2020 11:05:05",
          "content": "<p>How do you think,  should we trust cv or lb more? (Which of these models you`d choose as final score on private)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762238,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "03/03/2020 11:02:11",
      "content": "<p>Not huge I guess. If only there are really unuasual privat set words, like from old papers and libraries</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762373,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "03/03/2020 13:33:32",
      "content": "<p>My concern is where unseen graphemes are. If unseen graphemes are only in private test set, there would be a shake.\nI'm not sure how many unseen graphemes actually exist. The possible combinations are 12936 (168 * 11 * 7). There are 1292 graphemes in training set.\nThe <a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">slide</a> says that they selected 1295 commonly used bengali graphemes...</p>",
      "votes": null,
      "replies": [
        {
          "id": 762439,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "03/03/2020 14:08:35",
          "content": "<p>I guess all the unseen graphemes in the private LB. If they are split to both public and private, it is hard to achieve 0.99+. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 762620,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/03/2020 16:48:45",
      "content": "<p>no worry. do this test:</p>\n\n<ol>\n<li>train a classifier using 1295 grapheme class</li>\n<li>use this to decode into root, vowel, constant</li>\n<li>make a submission</li>\n</ol>\n\n<p>the cv/lb gap tells you  how many \"unseen graphemes\" there are ...\n(more correctly it tells you how many \"unidentifiable graphemes\"  there are )</p>\n\n<p>it think it is not a lot </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762961,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "03/04/2020 01:23:53",
      "content": "<p>Same for me, do observe some differences. I think this might be particularly alarming for such a dense LB. There can be 15 people in that range....\nLocal .9965: LB .9887\nLocal .9967; LB .9885</p>\n\n<p>Are those two experiments from the same model in your cases? My experiments above are from two (although very similar) models</p>",
      "votes": null,
      "replies": [
        {
          "id": 762991,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "03/04/2020 02:07:35",
          "content": "<p>yes same model ~</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "762213": "Two of my submissions:\n1) local 0.9968, LB 0.9884\n2) local 0.9972, LB 0.9881\nthe local - PB gap for my model (same model, same fold) ranges from 0.0083 to 0.0091\nthus I guess there may be ~0.0007 shake up for a single model (single fold)\n\n---Update---\nlocal 0.9978, LB 0.9895, my gap comes back into 0.0083",
    "762226": "glad you brought this topic. I'm also experiencing something similar. For the most part, CV and LB correlates well but there are times when local CV doesn't correlate well with LB; don't know why?\n```\nCV: 996409\nLB: 0.9889\n\nCV: 996608\nLB: 0.9883\n```\nEDIT: So yes, I also expect little shakeup :(",
    "762238": "Not huge I guess. If only there are really unuasual privat set words, like from old papers and libraries",
    "762373": "My concern is where unseen graphemes are. If unseen graphemes are only in private test set, there would be a shake.\nI'm not sure how many unseen graphemes actually exist. The possible combinations are 12936 (168 * 11 * 7). There are 1292 graphemes in training set.\nThe [slide](https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf) says that they selected 1295 commonly used bengali graphemes...",
    "762439": "I guess all the unseen graphemes in the private LB. If they are split to both public and private, it is hard to achieve 0.99+.",
    "762620": "no worry. do this test:\n\n1. train a classifier using 1295 grapheme class\n2. use this to decode into root, vowel, constant\n3. make a submission\n\nthe cv/lb gap tells you  how many \"unseen graphemes\" there are ...\n(more correctly it tells you how many \"unidentifiable graphemes\"  there are )\n\n it think it is not a lot",
    "762961": "Same for me, do observe some differences. I think this might be particularly alarming for such a dense LB. There can be 15 people in that range....\nLocal .9965: LB .9887\nLocal .9967; LB .9885\n\nAre those two experiments from the same model in your cases? My experiments above are from two (although very similar) models",
    "762985": "You have too much gap between CV and LB😂",
    "762991": "yes same model ~",
    "763000": "murphy89 I actually thought the same. But it seems the gap becomes greater the more your CV gets to 1.00😧 Had a small CV/LB gap myself of consistent ~.007",
    "764335": "How do you think,  should we trust cv or lb more? (Which of these models you`d choose as final score on private)"
  },
  "source": "meta"
}