{
  "id": 391203,
  "title": "CV vs LB - is everyone seeing a lower LB score?",
  "url": "/competitions/asl-signs/discussion/391203",
  "author_name": "Pietro Maldini",
  "post_date": "2023-02-28T17:26:13.375000",
  "votes": 13,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hi everyone.<br>\nI hope you are enjoying this challenge.</p>\n<p>I'm here for an important question. Is everyone facing a gap between CV and LB accuracy score?</p>\n<p>We may discuss here techniques to reduce this gap.</p>\n<p>My current best is 0.78 train accuracy, 0.67 validation accuracy and 0.48 LB accuracy.</p>\n<p>I tried using quantization aware training but found no benefit by trying it.</p>",
  "messages": [
    {
      "id": 2163552,
      "postDate": "2023-02-28T23:03:33.557Z",
      "content": "<p>Because Popsign has to be user independent, I like to do leave-one-signer-out cross validation for all 21-folds and then get an average and standard deviation.  That way I can do a t-test to determine if one method is going to be better than another. </p>",
      "rawMarkdown": "Because Popsign has to be user independent, I like to do leave-one-signer-out cross validation for all 21-folds and then get an average and standard deviation.  That way I can do a t-test to determine if one method is going to be better than another. ",
      "votes": 16,
      "replies": [
        {
          "id": 2164058,
          "postDate": "2023-03-01T09:23:10.037Z",
          "content": "<p>Does this mean that there are new participants in hidden test sets?</p>",
          "rawMarkdown": "Does this mean that there are new participants in hidden test sets?"
        }
      ]
    },
    {
      "id": 2163249,
      "postDate": "2023-02-28T17:26:13.377Z",
      "content": "<p>Hi everyone.<br>\nI hope you are enjoying this challenge.</p>\n<p>I'm here for an important question. Is everyone facing a gap between CV and LB accuracy score?</p>\n<p>We may discuss here techniques to reduce this gap.</p>\n<p>My current best is 0.78 train accuracy, 0.67 validation accuracy and 0.48 LB accuracy.</p>\n<p>I tried using quantization aware training but found no benefit by trying it.</p>",
      "rawMarkdown": "Hi everyone.\nI hope you are enjoying this challenge.\n\nI'm here for an important question. Is everyone facing a gap between CV and LB accuracy score?\n\nWe may discuss here techniques to reduce this gap.\n\nMy current best is 0.78 train accuracy, 0.67 validation accuracy and 0.48 LB accuracy.\n\nI tried using quantization aware training but found no benefit by trying it.",
      "votes": 13
    },
    {
      "id": 2163584,
      "postDate": "2023-02-28T23:38:22.087Z",
      "content": "<p>StratifiedgroupKFold: Stratified by sign grouped by participants</p>\n<p>CV 0.56<br>\nLB 0.60<br>\nbut much longer running time than expected </p>",
      "rawMarkdown": "StratifiedgroupKFold: Stratified by sign grouped by participants\n\nCV 0.56\nLB 0.60\nbut much longer running time than expected ",
      "votes": 7,
      "replies": [
        {
          "id": 2163604,
          "postDate": "2023-03-01T00:13:45.700Z",
          "content": "<p>Nicely done! Was your LB score an ensemble? That is, a model that internally ensembled all the KFold models?</p>",
          "rawMarkdown": "Nicely done! Was your LB score an ensemble? That is, a model that internally ensembled all the KFold models?",
          "replies": [
            {
              "id": 2163606,
              "postDate": "2023-03-01T00:19:28.223Z",
              "content": "<p>not yet, it is only one fold model</p>",
              "rawMarkdown": "not yet, it is only one fold model",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2169300,
      "postDate": "2023-03-05T02:08:56.530Z",
      "content": "<p>My new public notebook gets 0.74 optimistic CV, and 0.63 LB</p>\n<p><a href=\"https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders</a></p>",
      "rawMarkdown": "My new public notebook gets 0.74 optimistic CV, and 0.63 LB\n\nhttps://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders",
      "votes": 5
    },
    {
      "id": 2165407,
      "postDate": "2023-03-02T07:29:45.477Z",
      "content": "<p>same fold from this notebook -&gt; <a href=\"https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission\" target=\"_blank\">https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission</a><br>\n<code>train_test_split(datax, datay, test_size=0.15, random_state=42)</code></p>\n<p>CV 0.52 LB 0.45<br>\nCV 0.55 LB 0.48<br>\nCV 0.57 LB 0.49<br>\nCV 0.59 LB 0.50</p>\n<p>CV and LB correlate very well!</p>",
      "rawMarkdown": "same fold from this notebook -> https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission\n`train_test_split(datax, datay, test_size=0.15, random_state=42)`\n\nCV 0.52 LB 0.45\nCV 0.55 LB 0.48\nCV 0.57 LB 0.49\nCV 0.59 LB 0.50\n\nCV and LB correlate very well!",
      "votes": 3,
      "replies": [
        {
          "id": 2166813,
          "postDate": "2023-03-03T03:24:15.440Z",
          "content": "<p>I improved the model from here, but the gap between LB and CV increased tremendously.<br>\nProbably because of the leak. I need to groupKFold it with participant_id like Carno.</p>",
          "rawMarkdown": "I improved the model from here, but the gap between LB and CV increased tremendously.\nProbably because of the leak. I need to groupKFold it with participant_id like Carno.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2167450,
      "postDate": "2023-03-03T13:49:24.327Z",
      "content": "<p>Using StratifiedGroupKFold, I get the following results from some models using one fold with the following setup: <a href=\"https://www.kaggle.com/code/clemchris/asl-sign-detection-pytorch-lightning\" target=\"_blank\">https://www.kaggle.com/code/clemchris/asl-sign-detection-pytorch-lightning</a></p>\n<p>Fold 0, Val Acc: 0.4304, LB: 0.48<br>\nFold 2, Val Acc: 0.4700, LB: 0.52<br>\nFold 4, Val Acc: 0.4572, LB: 0.52</p>",
      "rawMarkdown": "Using StratifiedGroupKFold, I get the following results from some models using one fold with the following setup: https://www.kaggle.com/code/clemchris/asl-sign-detection-pytorch-lightning\n\nFold 0, Val Acc: 0.4304, LB: 0.48\nFold 2, Val Acc: 0.4700, LB: 0.52\nFold 4, Val Acc: 0.4572, LB: 0.52",
      "votes": 4,
      "replies": [
        {
          "id": 2168613,
          "postDate": "2023-03-04T11:24:59.630Z",
          "content": "<p><a href=\"https://www.kaggle.com/clemchris\" target=\"_blank\">@clemchris</a> </p>\n<p>thanks for reporting your results. are the results reported based on the same hyper-parameters in the notebook or after hyper-parameters optimization?</p>\n<p>i tried to used the same hyper-parameters but could not get the same results.</p>\n<hr>\n<p>by the way, you are using mean of the frames as input feature. if that works, it will also single frame (or the frame and neighbors that is close to mean) as input might work to some extent.<br>\nthe next step would probably see how you can pool better over the whole video (e.g. multiple instance, attention pool/average), rather than just average</p>",
          "rawMarkdown": "@clemchris \n\nthanks for reporting your results. are the results reported based on the same hyper-parameters in the notebook or after hyper-parameters optimization?\n\ni tried to used the same hyper-parameters but could not get the same results.\n\n---\n\nby the way, you are using mean of the frames as input feature. if that works, it will also single frame (or the frame and neighbors that is close to mean) as input might work to some extent.\nthe next step would probably see how you can pool better over the whole video (e.g. multiple instance, attention pool/average), rather than just average",
          "votes": 1
        },
        {
          "id": 2169391,
          "postDate": "2023-03-05T04:27:05.137Z",
          "content": "<p>Interesting that your LB is higher than the CV, I only saw people reporting lower LB than CV. Do you do something special that the others might not?</p>",
          "rawMarkdown": "Interesting that your LB is higher than the CV, I only saw people reporting lower LB than CV. Do you do something special that the others might not?",
          "votes": 1,
          "replies": [
            {
              "id": 2169794,
              "postDate": "2023-03-05T12:47:29.333Z",
              "content": "<p>So I guess most people don't split train and valid on a participant_id- but on a video level.</p>",
              "rawMarkdown": "So I guess most people don't split train and valid on a participant_id- but on a video level.",
              "votes": 1
            },
            {
              "id": 2170736,
              "postDate": "2023-03-06T08:51:07.023Z",
              "content": "<p>Yes, I think this might be the reason. I am still seeing it when using less features (lips instead of all face, see the latest version of the notebook)</p>",
              "rawMarkdown": "Yes, I think this might be the reason. I am still seeing it when using less features (lips instead of all face, see the latest version of the notebook)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2163506,
      "postDate": "2023-02-28T21:48:07.633Z",
      "content": "<p>it is better to do a few splitting and record results of:</p>\n<ul>\n<li>CV within user</li>\n<li>CV between user</li>\n<li>CV for each words for each of  words  …</li>\n</ul>\n<p>we need to visualise the data (tsne, etc) to see where the variation lies</p>",
      "rawMarkdown": "it is better to do a few splitting and record results of:\n- CV within user\n- CV between user\n- CV for each words for each of  words  ...\n\nwe need to visualise the data (tsne, etc) to see where the variation lies",
      "votes": 4,
      "replies": [
        {
          "id": 2163532,
          "postDate": "2023-02-28T22:28:20.763Z",
          "content": "<p>You are absolutely correct.</p>\n<p>I think also I should go back some steps and analyze difference between signs between sequences of the same user and between sequences of different users.</p>\n<p>Finding those differences may help finding ways to \"normalize\" the users  just from a sequence or also it may help in finding useful ways to augment the dataset without adding external data.</p>",
          "rawMarkdown": "You are absolutely correct.\n\nI think also I should go back some steps and analyze difference between signs between sequences of the same user and between sequences of different users.\n\nFinding those differences may help finding ways to \"normalize\" the users  just from a sequence or also it may help in finding useful ways to augment the dataset without adding external data."
        }
      ]
    },
    {
      "id": 2163320,
      "postDate": "2023-02-28T18:05:43Z",
      "content": "<p>Yes, I'm seeing similar gap between CV and LB scores.  I wonder whether it's related to overfitting to the individual signers in the training set.  Have you tried holding out all the samples from some of the signers and using those for validation?  (I haven't yet.)</p>",
      "rawMarkdown": "Yes, I'm seeing similar gap between CV and LB scores.  I wonder whether it's related to overfitting to the individual signers in the training set.  Have you tried holding out all the samples from some of the signers and using those for validation?  (I haven't yet.)",
      "votes": 1,
      "replies": [
        {
          "id": 2163364,
          "postDate": "2023-02-28T18:53:26.230Z",
          "content": "<p>Oh, yes, that's probably why. To get a more accurate CV, we will need to do like 7-fold GroupKFold, holding out 3 signers each fold. Note that will improve the correlation between CV and LB, but probably won't help LB any (even if you figure out how to have a single model that ensembles the base models?) It's mainly for getting more realistic CV.</p>",
          "rawMarkdown": "Oh, yes, that's probably why. To get a more accurate CV, we will need to do like 7-fold GroupKFold, holding out 3 signers each fold. Note that will improve the correlation between CV and LB, but probably won't help LB any (even if you figure out how to have a single model that ensembles the base models?) It's mainly for getting more realistic CV.",
          "votes": 1
        },
        {
          "id": 2163501,
          "postDate": "2023-02-28T21:35:58.163Z",
          "content": "<p>I tried right now splitting signers, I noticed that my CV dropped from 0.67 to around 0.35.<br>\nThe model seemed to be heavily overfitting to the individual signers. <br>\nI'll try to find some way to get better performance with this split and see if  with this kind of split I see a better correlation.</p>",
          "rawMarkdown": "I tried right now splitting signers, I noticed that my CV dropped from 0.67 to around 0.35.\nThe model seemed to be heavily overfitting to the individual signers. \nI'll try to find some way to get better performance with this split and see if  with this kind of split I see a better correlation.",
          "votes": 1,
          "replies": [
            {
              "id": 2163544,
              "postDate": "2023-02-28T22:50:34.927Z",
              "content": "<p>So, as I feared, it looks like the variability between signers is greater than the variability between a single signer doing a sign multiple times.  Not exactly surprising, but given we've only got 21 signers in the test set, that's going to make it a challenge.</p>",
              "rawMarkdown": "So, as I feared, it looks like the variability between signers is greater than the variability between a single signer doing a sign multiple times.  Not exactly surprising, but given we've only got 21 signers in the test set, that's going to make it a challenge."
            },
            {
              "id": 2163554,
              "postDate": "2023-02-28T23:04:16.383Z",
              "content": "<p>And of the 21, how many sign left-handed vs right-handed? Most likely it makes this issue even worse</p>",
              "rawMarkdown": "And of the 21, how many sign left-handed vs right-handed? Most likely it makes this issue even worse"
            },
            {
              "id": 2163562,
              "postDate": "2023-02-28T23:12:31.693Z",
              "content": "<p>One trick my students and I did to make some test models is just to sum the deltas of both hands.  That is a crude way to handle handedness.  Another potential option is to mirror all examples so you get the same number of left and right handed signs.  Offhand, I can not think of any signs in this vocabulary for which that will not work, but no guarantees. I'll have to try it myself.</p>",
              "rawMarkdown": "One trick my students and I did to make some test models is just to sum the deltas of both hands.  That is a crude way to handle handedness.  Another potential option is to mirror all examples so you get the same number of left and right handed signs.  Offhand, I can not think of any signs in this vocabulary for which that will not work, but no guarantees. I'll have to try it myself.",
              "votes": 4
            },
            {
              "id": 2163569,
              "postDate": "2023-02-28T23:18:32.607Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2166918,
      "postDate": "2023-03-03T05:48:45.893Z",
      "content": "<p>In my experiments, CV (single fold validation) &amp; LB correlate well too (usually there's 0.07 gap between the scores).</p>\n<p>CV 0.6310 LB 0.56<br>\nCV 0.6677 LB 0.60</p>",
      "rawMarkdown": "In my experiments, CV (single fold validation) & LB correlate well too (usually there's 0.07 gap between the scores).\n\nCV 0.6310 LB 0.56\nCV 0.6677 LB 0.60",
      "votes": 2,
      "replies": [
        {
          "id": 2168348,
          "postDate": "2023-03-04T06:32:33.283Z",
          "content": "<p>That's the same gap as what I got on my first working LB: CV 0.64, LB 0.57.</p>",
          "rawMarkdown": "That's the same gap as what I got on my first working LB: CV 0.64, LB 0.57.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2164085,
      "postDate": "2023-03-01T09:39:46.430Z",
      "content": "<p>In addition to the problem I mentioned below (with very few signers in the training set and apparently large variability between signers), I also wonder how much loss of accuracy there is when the models are converted to TFLite.  Presumably this would be easy enough to measure - just run your validation data through the final TFLite model - but my hunch is that the loss there will be very small.</p>",
      "rawMarkdown": "In addition to the problem I mentioned below (with very few signers in the training set and apparently large variability between signers), I also wonder how much loss of accuracy there is when the models are converted to TFLite.  Presumably this would be easy enough to measure - just run your validation data through the final TFLite model - but my hunch is that the loss there will be very small."
    },
    {
      "id": 2169885,
      "postDate": "2023-03-05T14:24:37.350Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2163552,
      "author_name": "Thad Starner",
      "author_url": "",
      "post_date": "2023-02-28T23:03:33.557000",
      "content": "<p>Because Popsign has to be user independent, I like to do leave-one-signer-out cross validation for all 21-folds and then get an average and standard deviation.  That way I can do a t-test to determine if one method is going to be better than another. </p>",
      "votes": 16,
      "replies": [
        {
          "id": 2164058,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2023-03-01T09:23:10.037000",
          "content": "<p>Does this mean that there are new participants in hidden test sets?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2163584,
      "author_name": "Carno Zhao",
      "author_url": "",
      "post_date": "2023-02-28T23:38:22.087000",
      "content": "<p>StratifiedgroupKFold: Stratified by sign grouped by participants</p>\n<p>CV 0.56<br>\nLB 0.60<br>\nbut much longer running time than expected </p>",
      "votes": 7,
      "replies": [
        {
          "id": 2163604,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2023-03-01T00:13:45.700000",
          "content": "<p>Nicely done! Was your LB score an ensemble? That is, a model that internally ensembled all the KFold models?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2163606,
              "author_name": "Carno Zhao",
              "author_url": "",
              "post_date": "2023-03-01T00:19:28.223000",
              "content": "<p>not yet, it is only one fold model</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2169300,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-03-05T02:08:56.530000",
      "content": "<p>My new public notebook gets 0.74 optimistic CV, and 0.63 LB</p>\n<p><a href=\"https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2165407,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2023-03-02T07:29:45.477000",
      "content": "<p>same fold from this notebook -&gt; <a href=\"https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission\" target=\"_blank\">https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission</a><br>\n<code>train_test_split(datax, datay, test_size=0.15, random_state=42)</code></p>\n<p>CV 0.52 LB 0.45<br>\nCV 0.55 LB 0.48<br>\nCV 0.57 LB 0.49<br>\nCV 0.59 LB 0.50</p>\n<p>CV and LB correlate very well!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2166813,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2023-03-03T03:24:15.440000",
          "content": "<p>I improved the model from here, but the gap between LB and CV increased tremendously.<br>\nProbably because of the leak. I need to groupKFold it with participant_id like Carno.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2167450,
      "author_name": "clem-chris",
      "author_url": "",
      "post_date": "2023-03-03T13:49:24.327000",
      "content": "<p>Using StratifiedGroupKFold, I get the following results from some models using one fold with the following setup: <a href=\"https://www.kaggle.com/code/clemchris/asl-sign-detection-pytorch-lightning\" target=\"_blank\">https://www.kaggle.com/code/clemchris/asl-sign-detection-pytorch-lightning</a></p>\n<p>Fold 0, Val Acc: 0.4304, LB: 0.48<br>\nFold 2, Val Acc: 0.4700, LB: 0.52<br>\nFold 4, Val Acc: 0.4572, LB: 0.52</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2168613,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-04T11:24:59.630000",
          "content": "<p><a href=\"https://www.kaggle.com/clemchris\" target=\"_blank\">@clemchris</a> </p>\n<p>thanks for reporting your results. are the results reported based on the same hyper-parameters in the notebook or after hyper-parameters optimization?</p>\n<p>i tried to used the same hyper-parameters but could not get the same results.</p>\n<hr>\n<p>by the way, you are using mean of the frames as input feature. if that works, it will also single frame (or the frame and neighbors that is close to mean) as input might work to some extent.<br>\nthe next step would probably see how you can pool better over the whole video (e.g. multiple instance, attention pool/average), rather than just average</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2169391,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2023-03-05T04:27:05.137000",
          "content": "<p>Interesting that your LB is higher than the CV, I only saw people reporting lower LB than CV. Do you do something special that the others might not?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2169794,
              "author_name": "nofreewill42",
              "author_url": "",
              "post_date": "2023-03-05T12:47:29.333000",
              "content": "<p>So I guess most people don't split train and valid on a participant_id- but on a video level.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2170736,
              "author_name": "clem-chris",
              "author_url": "",
              "post_date": "2023-03-06T08:51:07.023000",
              "content": "<p>Yes, I think this might be the reason. I am still seeing it when using less features (lips instead of all face, see the latest version of the notebook)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2163506,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-28T21:48:07.633000",
      "content": "<p>it is better to do a few splitting and record results of:</p>\n<ul>\n<li>CV within user</li>\n<li>CV between user</li>\n<li>CV for each words for each of  words  …</li>\n</ul>\n<p>we need to visualise the data (tsne, etc) to see where the variation lies</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2163532,
          "author_name": "Pietro Maldini",
          "author_url": "",
          "post_date": "2023-02-28T22:28:20.763000",
          "content": "<p>You are absolutely correct.</p>\n<p>I think also I should go back some steps and analyze difference between signs between sequences of the same user and between sequences of different users.</p>\n<p>Finding those differences may help finding ways to \"normalize\" the users  just from a sequence or also it may help in finding useful ways to augment the dataset without adding external data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2163320,
      "author_name": "Andrew",
      "author_url": "",
      "post_date": "2023-02-28T18:05:43",
      "content": "<p>Yes, I'm seeing similar gap between CV and LB scores.  I wonder whether it's related to overfitting to the individual signers in the training set.  Have you tried holding out all the samples from some of the signers and using those for validation?  (I haven't yet.)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2163364,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2023-02-28T18:53:26.230000",
          "content": "<p>Oh, yes, that's probably why. To get a more accurate CV, we will need to do like 7-fold GroupKFold, holding out 3 signers each fold. Note that will improve the correlation between CV and LB, but probably won't help LB any (even if you figure out how to have a single model that ensembles the base models?) It's mainly for getting more realistic CV.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2163501,
          "author_name": "Pietro Maldini",
          "author_url": "",
          "post_date": "2023-02-28T21:35:58.163000",
          "content": "<p>I tried right now splitting signers, I noticed that my CV dropped from 0.67 to around 0.35.<br>\nThe model seemed to be heavily overfitting to the individual signers. <br>\nI'll try to find some way to get better performance with this split and see if  with this kind of split I see a better correlation.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2163544,
              "author_name": "Andrew",
              "author_url": "",
              "post_date": "2023-02-28T22:50:34.927000",
              "content": "<p>So, as I feared, it looks like the variability between signers is greater than the variability between a single signer doing a sign multiple times.  Not exactly surprising, but given we've only got 21 signers in the test set, that's going to make it a challenge.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2163554,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-28T23:04:16.383000",
              "content": "<p>And of the 21, how many sign left-handed vs right-handed? Most likely it makes this issue even worse</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2163562,
              "author_name": "Thad Starner",
              "author_url": "",
              "post_date": "2023-02-28T23:12:31.693000",
              "content": "<p>One trick my students and I did to make some test models is just to sum the deltas of both hands.  That is a crude way to handle handedness.  Another potential option is to mirror all examples so you get the same number of left and right handed signs.  Offhand, I can not think of any signs in this vocabulary for which that will not work, but no guarantees. I'll have to try it myself.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2163569,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-28T23:18:32.607000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2166918,
      "author_name": "HyeongChan Kim",
      "author_url": "",
      "post_date": "2023-03-03T05:48:45.893000",
      "content": "<p>In my experiments, CV (single fold validation) &amp; LB correlate well too (usually there's 0.07 gap between the scores).</p>\n<p>CV 0.6310 LB 0.56<br>\nCV 0.6677 LB 0.60</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2168348,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2023-03-04T06:32:33.283000",
          "content": "<p>That's the same gap as what I got on my first working LB: CV 0.64, LB 0.57.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2164085,
      "author_name": "Andrew",
      "author_url": "",
      "post_date": "2023-03-01T09:39:46.430000",
      "content": "<p>In addition to the problem I mentioned below (with very few signers in the training set and apparently large variability between signers), I also wonder how much loss of accuracy there is when the models are converted to TFLite.  Presumably this would be easy enough to measure - just run your validation data through the final TFLite model - but my hunch is that the loss there will be very small.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2169885,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-05T14:24:37.350000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2163552": "Because Popsign has to be user independent, I like to do leave-one-signer-out cross validation for all 21-folds and then get an average and standard deviation.  That way I can do a t-test to determine if one method is going to be better than another. ",
    "2163249": "Hi everyone.\nI hope you are enjoying this challenge.\n\nI'm here for an important question. Is everyone facing a gap between CV and LB accuracy score?\n\nWe may discuss here techniques to reduce this gap.\n\nMy current best is 0.78 train accuracy, 0.67 validation accuracy and 0.48 LB accuracy.\n\nI tried using quantization aware training but found no benefit by trying it.",
    "2163584": "StratifiedgroupKFold: Stratified by sign grouped by participants\n\nCV 0.56\nLB 0.60\nbut much longer running time than expected ",
    "2169300": "My new public notebook gets 0.74 optimistic CV, and 0.63 LB\n\nhttps://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders",
    "2165407": "same fold from this notebook -> https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission\n`train_test_split(datax, datay, test_size=0.15, random_state=42)`\n\nCV 0.52 LB 0.45\nCV 0.55 LB 0.48\nCV 0.57 LB 0.49\nCV 0.59 LB 0.50\n\nCV and LB correlate very well!",
    "2167450": "Using StratifiedGroupKFold, I get the following results from some models using one fold with the following setup: https://www.kaggle.com/code/clemchris/asl-sign-detection-pytorch-lightning\n\nFold 0, Val Acc: 0.4304, LB: 0.48\nFold 2, Val Acc: 0.4700, LB: 0.52\nFold 4, Val Acc: 0.4572, LB: 0.52",
    "2163506": "it is better to do a few splitting and record results of:\n- CV within user\n- CV between user\n- CV for each words for each of  words  ...\n\nwe need to visualise the data (tsne, etc) to see where the variation lies",
    "2163320": "Yes, I'm seeing similar gap between CV and LB scores.  I wonder whether it's related to overfitting to the individual signers in the training set.  Have you tried holding out all the samples from some of the signers and using those for validation?  (I haven't yet.)",
    "2166918": "In my experiments, CV (single fold validation) & LB correlate well too (usually there's 0.07 gap between the scores).\n\nCV 0.6310 LB 0.56\nCV 0.6677 LB 0.60",
    "2164085": "In addition to the problem I mentioned below (with very few signers in the training set and apparently large variability between signers), I also wonder how much loss of accuracy there is when the models are converted to TFLite.  Presumably this would be easy enough to measure - just run your validation data through the final TFLite model - but my hunch is that the loss there will be very small.",
    "2169885": ""
  }
}