{
  "id": 406364,
  "title": "Why private test score is so high?",
  "url": "/competitions/asl-signs/discussion/406364",
  "author_name": "",
  "post_date": "2023-05-02T04:46:59.020895800Z",
  "votes": 14,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Gold zone score is above 0.87.<br>\nI don't think it can be achived by participant split, at leaset for training set.</p>\n<p>Does private test set have many same participants with training set?<br>\nOr, is there any other reason?</p>",
  "messages": [
    {
      "id": "2242164",
      "postDate": "05/02/2023 04:46:59",
      "content": "<p>Gold zone score is above 0.87.<br>\nI don't think it can be achived by participant split, at leaset for training set.</p>\n<p>Does private test set have many same participants with training set?<br>\nOr, is there any other reason?</p>",
      "rawMarkdown": "Gold zone score is above 0.87.\nI don't think it can be achived by participant split, at leaset for training set.\n\nDoes private test set have many same participants with training set?\nOr, is there any other reason?",
      "votes": null
    },
    {
      "id": "2242189",
      "postDate": "05/02/2023 05:16:40",
      "content": "<p>I have the same doubt. It seems for every model, the accuracy has increased for the private lb. Even the models that didn't perform well at all on public LB. </p>",
      "rawMarkdown": "I have the same doubt. It seems for every model, the accuracy has increased for the private lb. Even the models that didn't perform well at all on public LB.",
      "votes": null
    },
    {
      "id": "2242228",
      "postDate": "05/02/2023 06:11:10",
      "content": "<p>I agree it is very weird. The public LB is similar to local CV score but the private LB is not.</p>",
      "rawMarkdown": "I agree it is very weird. The public LB is similar to local CV score but the private LB is not.",
      "votes": null
    },
    {
      "id": "2242333",
      "postDate": "05/02/2023 07:35:57",
      "content": "<p>They were very similar to my private CV scores, 0.01-0.02 higher than my local scores.  I was surprised my public LB scores were so low. </p>",
      "rawMarkdown": "They were very similar to my private CV scores, 0.01-0.02 higher than my local scores.  I was surprised my public LB scores were so low.",
      "votes": null
    },
    {
      "id": "2242430",
      "postDate": "05/02/2023 09:16:42",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <br>\nCan you follow up on that ? </p>\n<p>On my local validation setup, I have only one patient with CV &gt; 0.87 (id=26734)</p>\n<p>The private LB scores do not look right.<br>\nEither there is train / test overlap or the private test is surprisingly not noisy.</p>",
      "rawMarkdown": "sohier @ashleychow \nCan you follow up on that ? \n\nOn my local validation setup, I have only one patient with CV > 0.87 (id=26734)\n\nThe private LB scores do not look right.\nEither there is train / test overlap or the private test is surprisingly not noisy.",
      "votes": null
    },
    {
      "id": "2242444",
      "postDate": "05/02/2023 09:36:01",
      "content": "<p>For reference, using a single model: </p>\n<ul>\n<li>Validation CV : 0.74</li>\n<li>Training CV : 0.94</li>\n<li>Public LB : 0.78, private LB 0.87  (not verified)</li>\n</ul>",
      "rawMarkdown": "For reference, using a single model: \n- Validation CV : 0.74\n- Training CV : 0.94\n- Public LB : 0.78, private LB 0.87  (not verified)",
      "votes": null
    },
    {
      "id": "2242468",
      "postDate": "05/02/2023 09:56:42",
      "content": "<p>Thanks for sharing Theo. Also congrats for your result!</p>\n<p>May I Ask how you split training data? I guess it's split by participant, but just making sure:)</p>",
      "rawMarkdown": "Thanks for sharing Theo. Also congrats for your result!\n\nMay I Ask how you split training data? I guess it's split by participant, but just making sure:)",
      "votes": null
    },
    {
      "id": "2242481",
      "postDate": "05/02/2023 10:12:38",
      "content": "<p>I think they used non-overlapping participant on the public, and the overlapping on the private.</p>\n<p>My single model CV with participant split:~0.80, public LB:~0.80<br>\nrandom split CV:~0.88(not verified), private LB:~0.88</p>\n<p>random split CV is not verified but very likely(I did not check the random split CV after I crossed 0.87)</p>",
      "rawMarkdown": "I think they used non-overlapping participant on the public, and the overlapping on the private.\n\nMy single model CV with participant split:~0.80, public LB:~0.80\nrandom split CV:~0.88(not verified), private LB:~0.88\n\nrandom split CV is not verified but very likely(I did not check the random split CV after I crossed 0.87)",
      "votes": null
    },
    {
      "id": "2242527",
      "postDate": "05/02/2023 10:40:22",
      "content": "<p>Thanks and congratz to you as well !</p>\n<p>I split by participant yes (groupkfold with k=4), as recommended by the hosts.</p>",
      "rawMarkdown": "Thanks and congratz to you as well !\n\nI split by participant yes (groupkfold with k=4), as recommended by the hosts.",
      "votes": null
    },
    {
      "id": "2242536",
      "postDate": "05/02/2023 10:46:58",
      "content": "<p>That is also my guess. My CV scores perfectly align in both cases, too.</p>",
      "rawMarkdown": "That is also my guess. My CV scores perfectly align in both cases, too.",
      "votes": null
    },
    {
      "id": "2242637",
      "postDate": "05/02/2023 12:16:59",
      "content": "<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> Thanks for sharing, congrats for the solo winning! Looking forward your solution by the way;)</p>",
      "rawMarkdown": "hoyso48 Thanks for sharing, congrats for the solo winning! Looking forward your solution by the way;)",
      "votes": null
    },
    {
      "id": "2242671",
      "postDate": "05/02/2023 12:43:58",
      "content": "<p>private set is cleaned data?</p>",
      "rawMarkdown": "private set is cleaned data?",
      "votes": null
    },
    {
      "id": "2243140",
      "postDate": "05/02/2023 17:47:56",
      "content": "<p>I tried overfitting a model on the training data, it performed worse.<br>\nSo I don't think there's duplicates between train and test. There can still be signer overlap though.</p>",
      "rawMarkdown": "I tried overfitting a model on the training data, it performed worse.\nSo I don't think there's duplicates between train and test. There can still be signer overlap though.",
      "votes": null
    },
    {
      "id": "2243173",
      "postDate": "05/02/2023 18:03:27",
      "content": "<p>haha, i'm trying the same experiment now too.</p>\n<p>If the private test data has signer (i.e. participant) overlap (but not duplicate videos), I wonder how we could design our models differently to take advantage of this?</p>",
      "rawMarkdown": "haha, i'm trying the same experiment now too.\n\nIf the private test data has signer (i.e. participant) overlap (but not duplicate videos), I wonder how we could design our models differently to take advantage of this?",
      "votes": null
    },
    {
      "id": "2243176",
      "postDate": "05/02/2023 18:04:22",
      "content": "<p>Does each participant only sign each word once? If so, perhaps we can classify the participant then remove any words they already signed in train data.</p>",
      "rawMarkdown": "Does each participant only sign each word once? If so, perhaps we can classify the participant then remove any words they already signed in train data.",
      "votes": null
    },
    {
      "id": "2243351",
      "postDate": "05/02/2023 20:55:08",
      "content": "<p>Now that the competition is over I can disclose that we split the dataset by signer with no overlap between the train, public, and private sets. There were only a handful of signers in the public set so we weren't too surprised to see that the public scores are somewhat idiosyncratic. </p>\n<p>Similarly, we expected to see a meaningful degree of variation in the difficulty of parsing videos from each signer. You can see some of that variation in the train set landmarks. Looking at the raw video reveals additional sources of difference, like lighting quality or distance from the camera.</p>\n<p>Overall I am not concerned.</p>",
      "rawMarkdown": "Now that the competition is over I can disclose that we split the dataset by signer with no overlap between the train, public, and private sets. There were only a handful of signers in the public set so we weren't too surprised to see that the public scores are somewhat idiosyncratic. \n\nSimilarly, we expected to see a meaningful degree of variation in the difficulty of parsing videos from each signer. You can see some of that variation in the train set landmarks. Looking at the raw video reveals additional sources of difference, like lighting quality or distance from the camera.\n\nOverall I am not concerned.",
      "votes": null
    },
    {
      "id": "2243393",
      "postDate": "05/02/2023 21:50:40",
      "content": "<p>Good to know, thanks Sohier !</p>",
      "rawMarkdown": "Good to know, thanks Sohier !",
      "votes": null
    },
    {
      "id": "2243422",
      "postDate": "05/02/2023 22:30:30",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for the clarification, good to know what really happened!</p>",
      "rawMarkdown": "Thanks @sohier for the clarification, good to know what really happened!",
      "votes": null
    },
    {
      "id": "2243430",
      "postDate": "05/02/2023 22:35:17",
      "content": "<p>I also trained my submission for 300 epochs instead of 120 and submitted. It's LB performed worse.</p>",
      "rawMarkdown": "I also trained my submission for 300 epochs instead of 120 and submitted. It's LB performed worse.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242189,
      "author_name": "ahsuna123",
      "author_url": "",
      "post_date": "05/02/2023 05:16:40",
      "content": "<p>I have the same doubt. It seems for every model, the accuracy has increased for the private lb. Even the models that didn't perform well at all on public LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2242228,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/02/2023 06:11:10",
      "content": "<p>I agree it is very weird. The public LB is similar to local CV score but the private LB is not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2242333,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "05/02/2023 07:35:57",
      "content": "<p>They were very similar to my private CV scores, 0.01-0.02 higher than my local scores.  I was surprised my public LB scores were so low. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2242430,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "05/02/2023 09:16:42",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/ashleychow\" target=\"_blank\">@ashleychow</a> <br>\nCan you follow up on that ? </p>\n<p>On my local validation setup, I have only one patient with CV &gt; 0.87 (id=26734)</p>\n<p>The private LB scores do not look right.<br>\nEither there is train / test overlap or the private test is surprisingly not noisy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2242444,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/02/2023 09:36:01",
          "content": "<p>For reference, using a single model: </p>\n<ul>\n<li>Validation CV : 0.74</li>\n<li>Training CV : 0.94</li>\n<li>Public LB : 0.78, private LB 0.87  (not verified)</li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2242468,
              "author_name": "bamps53",
              "author_url": "",
              "post_date": "05/02/2023 09:56:42",
              "content": "<p>Thanks for sharing Theo. Also congrats for your result!</p>\n<p>May I Ask how you split training data? I guess it's split by participant, but just making sure:)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2242527,
                  "author_name": "theoviel",
                  "author_url": "",
                  "post_date": "05/02/2023 10:40:22",
                  "content": "<p>Thanks and congratz to you as well !</p>\n<p>I split by participant yes (groupkfold with k=4), as recommended by the hosts.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2242481,
          "author_name": "hoyso48",
          "author_url": "",
          "post_date": "05/02/2023 10:12:38",
          "content": "<p>I think they used non-overlapping participant on the public, and the overlapping on the private.</p>\n<p>My single model CV with participant split:~0.80, public LB:~0.80<br>\nrandom split CV:~0.88(not verified), private LB:~0.88</p>\n<p>random split CV is not verified but very likely(I did not check the random split CV after I crossed 0.87)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2242536,
              "author_name": "gregorlied",
              "author_url": "",
              "post_date": "05/02/2023 10:46:58",
              "content": "<p>That is also my guess. My CV scores perfectly align in both cases, too.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2242637,
              "author_name": "bamps53",
              "author_url": "",
              "post_date": "05/02/2023 12:16:59",
              "content": "<p><a href=\"https://www.kaggle.com/hoyso48\" target=\"_blank\">@hoyso48</a> Thanks for sharing, congrats for the solo winning! Looking forward your solution by the way;)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242671,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/02/2023 12:43:58",
      "content": "<p>private set is cleaned data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2243140,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "05/02/2023 17:47:56",
      "content": "<p>I tried overfitting a model on the training data, it performed worse.<br>\nSo I don't think there's duplicates between train and test. There can still be signer overlap though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2243173,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/02/2023 18:03:27",
          "content": "<p>haha, i'm trying the same experiment now too.</p>\n<p>If the private test data has signer (i.e. participant) overlap (but not duplicate videos), I wonder how we could design our models differently to take advantage of this?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2243176,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "05/02/2023 18:04:22",
              "content": "<p>Does each participant only sign each word once? If so, perhaps we can classify the participant then remove any words they already signed in train data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2243430,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/02/2023 22:35:17",
          "content": "<p>I also trained my submission for 300 epochs instead of 120 and submitted. It's LB performed worse.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2243351,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "05/02/2023 20:55:08",
      "content": "<p>Now that the competition is over I can disclose that we split the dataset by signer with no overlap between the train, public, and private sets. There were only a handful of signers in the public set so we weren't too surprised to see that the public scores are somewhat idiosyncratic. </p>\n<p>Similarly, we expected to see a meaningful degree of variation in the difficulty of parsing videos from each signer. You can see some of that variation in the train set landmarks. Looking at the raw video reveals additional sources of difference, like lighting quality or distance from the camera.</p>\n<p>Overall I am not concerned.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2243393,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/02/2023 21:50:40",
          "content": "<p>Good to know, thanks Sohier !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2243422,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "05/02/2023 22:30:30",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for the clarification, good to know what really happened!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2242164": "Gold zone score is above 0.87.\nI don't think it can be achived by participant split, at leaset for training set.\n\nDoes private test set have many same participants with training set?\nOr, is there any other reason?",
    "2242189": "I have the same doubt. It seems for every model, the accuracy has increased for the private lb. Even the models that didn't perform well at all on public LB.",
    "2242228": "I agree it is very weird. The public LB is similar to local CV score but the private LB is not.",
    "2242333": "They were very similar to my private CV scores, 0.01-0.02 higher than my local scores.  I was surprised my public LB scores were so low.",
    "2242430": "sohier @ashleychow \nCan you follow up on that ? \n\nOn my local validation setup, I have only one patient with CV > 0.87 (id=26734)\n\nThe private LB scores do not look right.\nEither there is train / test overlap or the private test is surprisingly not noisy.",
    "2242444": "For reference, using a single model: \n- Validation CV : 0.74\n- Training CV : 0.94\n- Public LB : 0.78, private LB 0.87  (not verified)",
    "2242468": "Thanks for sharing Theo. Also congrats for your result!\n\nMay I Ask how you split training data? I guess it's split by participant, but just making sure:)",
    "2242481": "I think they used non-overlapping participant on the public, and the overlapping on the private.\n\nMy single model CV with participant split:~0.80, public LB:~0.80\nrandom split CV:~0.88(not verified), private LB:~0.88\n\nrandom split CV is not verified but very likely(I did not check the random split CV after I crossed 0.87)",
    "2242527": "Thanks and congratz to you as well !\n\nI split by participant yes (groupkfold with k=4), as recommended by the hosts.",
    "2242536": "That is also my guess. My CV scores perfectly align in both cases, too.",
    "2242637": "hoyso48 Thanks for sharing, congrats for the solo winning! Looking forward your solution by the way;)",
    "2242671": "private set is cleaned data?",
    "2243140": "I tried overfitting a model on the training data, it performed worse.\nSo I don't think there's duplicates between train and test. There can still be signer overlap though.",
    "2243173": "haha, i'm trying the same experiment now too.\n\nIf the private test data has signer (i.e. participant) overlap (but not duplicate videos), I wonder how we could design our models differently to take advantage of this?",
    "2243176": "Does each participant only sign each word once? If so, perhaps we can classify the participant then remove any words they already signed in train data.",
    "2243351": "Now that the competition is over I can disclose that we split the dataset by signer with no overlap between the train, public, and private sets. There were only a handful of signers in the public set so we weren't too surprised to see that the public scores are somewhat idiosyncratic. \n\nSimilarly, we expected to see a meaningful degree of variation in the difficulty of parsing videos from each signer. You can see some of that variation in the train set landmarks. Looking at the raw video reveals additional sources of difference, like lighting quality or distance from the camera.\n\nOverall I am not concerned.",
    "2243393": "Good to know, thanks Sohier !",
    "2243422": "Thanks @sohier for the clarification, good to know what really happened!",
    "2243430": "I also trained my submission for 300 epochs instead of 120 and submitted. It's LB performed worse."
  },
  "source": "meta"
}