{
  "id": 401883,
  "title": "CV vs Leaderboard",
  "url": "/competitions/asl-signs/discussion/401883",
  "author_name": "chemdatafarmer",
  "post_date": "2023-04-15T13:46:18.988000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey Everyone,</p>\n<p>I'm a bit curious how everyone's internal testing is lining up to the leaderboard. When I was getting a local CV of about 0.7, I was getting leaderboard scores of 0.65. Now I'm getting local CV of about 0.8 and LB scores are coming in at best around 0.7.</p>\n<p>I don't <em>think</em> I'm overfitting too badly because my train &amp; validation losses are essentially converging (validation is 10k random samples). </p>\n<p>Also, perhaps my random validation is too easy? I'm consistently getting a higher val accuracy and lower val loss up until the train/validation loss/accuracy converges. Anyone else see this before? I'm using relatively heavy dropout so I've been attributing it to that, but would be happy to hear other folks experiences/perspectives.</p>",
  "messages": [
    {
      "id": 2222865,
      "postDate": "2023-04-15T16:11:57.923Z",
      "content": "<blockquote>\n  <p>Also, perhaps my random validation is too easy?</p>\n</blockquote>\n<p>Yes, your validation leads to <a href=\"https://en.wikipedia.org/wiki/Leakage_(machine_learning)\" target=\"_blank\">data leakage</a>.<br>\nWe should do inference for unseen participants in test dataset, but your validation is not suitable for the situation.</p>\n<p>you should check previous discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391203\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391203</a></p>",
      "rawMarkdown": ">Also, perhaps my random validation is too easy?\n\nYes, your validation leads to [data leakage](https://en.wikipedia.org/wiki/Leakage_(machine_learning)).\nWe should do inference for unseen participants in test dataset, but your validation is not suitable for the situation.\n\nyou should check previous discussion.\n[https://www.kaggle.com/competitions/asl-signs/discussion/391203](https://www.kaggle.com/competitions/asl-signs/discussion/391203)",
      "votes": 1,
      "replies": [
        {
          "id": 2222887,
          "postDate": "2023-04-15T16:25:59.937Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a> thank you very much, this is very helpful. I had missed this previous discussion, thank you for bringing it to my attention! I will try with a participant_id split now and see how things go :)</p>",
          "rawMarkdown": "Hi @clearwaterkzk thank you very much, this is very helpful. I had missed this previous discussion, thank you for bringing it to my attention! I will try with a participant_id split now and see how things go :)"
        }
      ]
    },
    {
      "id": 2222816,
      "postDate": "2023-04-15T15:27:22.073Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F332d6d00212451f66a7ae8f086402ff2%2FModel_Accuracy_ASL_04152023%20-%20Copy.png?generation=1681572370772784&amp;alt=media\" alt=\"\"></p>\n<p>The above is the accuracy plots for a model I recently tried. Validation set was 79% accurate, leaderboard was 0.69.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F1271d9aeba43735b2964ed4b90683a9b%2FModel_Loss_ASL_04152023%20-%20Copy.png?generation=1681572380369029&amp;alt=media\" alt=\"\"></p>\n<p>The above are the loss plots for this same model.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F332d6d00212451f66a7ae8f086402ff2%2FModel_Accuracy_ASL_04152023%20-%20Copy.png?generation=1681572370772784&alt=media)\n\nThe above is the accuracy plots for a model I recently tried. Validation set was 79% accurate, leaderboard was 0.69.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F1271d9aeba43735b2964ed4b90683a9b%2FModel_Loss_ASL_04152023%20-%20Copy.png?generation=1681572380369029&alt=media)\n\nThe above are the loss plots for this same model.",
      "replies": [
        {
          "id": 2223579,
          "postDate": "2023-04-16T12:12:22.983Z",
          "content": "<p>Are you finding more success in cleaning the data or changing model architecture? </p>",
          "rawMarkdown": "Are you finding more success in cleaning the data or changing model architecture? ",
          "replies": [
            {
              "id": 2223594,
              "postDate": "2023-04-16T12:43:31.393Z",
              "content": "<p>Finding better ways to present the data to the model has had the biggest impacts on my scores so far. I've tried quite a few model architectures and so far attention based networks or conv nets (the above is a mixture of the two) perform somewhat similarly in my experiments, though there are public notebooks for \"transformers\" that do better than my own tests (or even my best models).</p>",
              "rawMarkdown": "Finding better ways to present the data to the model has had the biggest impacts on my scores so far. I've tried quite a few model architectures and so far attention based networks or conv nets (the above is a mixture of the two) perform somewhat similarly in my experiments, though there are public notebooks for \"transformers\" that do better than my own tests (or even my best models).",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2222700,
      "postDate": "2023-04-15T13:46:18.990Z",
      "content": "<p>Hey Everyone,</p>\n<p>I'm a bit curious how everyone's internal testing is lining up to the leaderboard. When I was getting a local CV of about 0.7, I was getting leaderboard scores of 0.65. Now I'm getting local CV of about 0.8 and LB scores are coming in at best around 0.7.</p>\n<p>I don't <em>think</em> I'm overfitting too badly because my train &amp; validation losses are essentially converging (validation is 10k random samples). </p>\n<p>Also, perhaps my random validation is too easy? I'm consistently getting a higher val accuracy and lower val loss up until the train/validation loss/accuracy converges. Anyone else see this before? I'm using relatively heavy dropout so I've been attributing it to that, but would be happy to hear other folks experiences/perspectives.</p>",
      "rawMarkdown": "Hey Everyone,\n\nI'm a bit curious how everyone's internal testing is lining up to the leaderboard. When I was getting a local CV of about 0.7, I was getting leaderboard scores of 0.65. Now I'm getting local CV of about 0.8 and LB scores are coming in at best around 0.7.\n\nI don't *think* I'm overfitting too badly because my train & validation losses are essentially converging (validation is 10k random samples). \n\nAlso, perhaps my random validation is too easy? I'm consistently getting a higher val accuracy and lower val loss up until the train/validation loss/accuracy converges. Anyone else see this before? I'm using relatively heavy dropout so I've been attributing it to that, but would be happy to hear other folks experiences/perspectives."
    },
    {
      "id": 2222812,
      "postDate": "2023-04-15T15:25:44.710Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2222865,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2023-04-15T16:11:57.923000",
      "content": "<blockquote>\n  <p>Also, perhaps my random validation is too easy?</p>\n</blockquote>\n<p>Yes, your validation leads to <a href=\"https://en.wikipedia.org/wiki/Leakage_(machine_learning)\" target=\"_blank\">data leakage</a>.<br>\nWe should do inference for unseen participants in test dataset, but your validation is not suitable for the situation.</p>\n<p>you should check previous discussion.<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391203\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391203</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2222887,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-04-15T16:25:59.937000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a> thank you very much, this is very helpful. I had missed this previous discussion, thank you for bringing it to my attention! I will try with a participant_id split now and see how things go :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2222816,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2023-04-15T15:27:22.073000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F332d6d00212451f66a7ae8f086402ff2%2FModel_Accuracy_ASL_04152023%20-%20Copy.png?generation=1681572370772784&amp;alt=media\" alt=\"\"></p>\n<p>The above is the accuracy plots for a model I recently tried. Validation set was 79% accurate, leaderboard was 0.69.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F1271d9aeba43735b2964ed4b90683a9b%2FModel_Loss_ASL_04152023%20-%20Copy.png?generation=1681572380369029&amp;alt=media\" alt=\"\"></p>\n<p>The above are the loss plots for this same model.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2223579,
          "author_name": "AbdouE",
          "author_url": "",
          "post_date": "2023-04-16T12:12:22.983000",
          "content": "<p>Are you finding more success in cleaning the data or changing model architecture? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2223594,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2023-04-16T12:43:31.393000",
              "content": "<p>Finding better ways to present the data to the model has had the biggest impacts on my scores so far. I've tried quite a few model architectures and so far attention based networks or conv nets (the above is a mixture of the two) perform somewhat similarly in my experiments, though there are public notebooks for \"transformers\" that do better than my own tests (or even my best models).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2222812,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-15T15:25:44.710000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2222865": ">Also, perhaps my random validation is too easy?\n\nYes, your validation leads to [data leakage](https://en.wikipedia.org/wiki/Leakage_(machine_learning)).\nWe should do inference for unseen participants in test dataset, but your validation is not suitable for the situation.\n\nyou should check previous discussion.\n[https://www.kaggle.com/competitions/asl-signs/discussion/391203](https://www.kaggle.com/competitions/asl-signs/discussion/391203)",
    "2222816": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F332d6d00212451f66a7ae8f086402ff2%2FModel_Accuracy_ASL_04152023%20-%20Copy.png?generation=1681572370772784&alt=media)\n\nThe above is the accuracy plots for a model I recently tried. Validation set was 79% accurate, leaderboard was 0.69.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1921811%2F1271d9aeba43735b2964ed4b90683a9b%2FModel_Loss_ASL_04152023%20-%20Copy.png?generation=1681572380369029&alt=media)\n\nThe above are the loss plots for this same model.",
    "2222700": "Hey Everyone,\n\nI'm a bit curious how everyone's internal testing is lining up to the leaderboard. When I was getting a local CV of about 0.7, I was getting leaderboard scores of 0.65. Now I'm getting local CV of about 0.8 and LB scores are coming in at best around 0.7.\n\nI don't *think* I'm overfitting too badly because my train & validation losses are essentially converging (validation is 10k random samples). \n\nAlso, perhaps my random validation is too easy? I'm consistently getting a higher val accuracy and lower val loss up until the train/validation loss/accuracy converges. Anyone else see this before? I'm using relatively heavy dropout so I've been attributing it to that, but would be happy to hear other folks experiences/perspectives.",
    "2222812": ""
  }
}