{
  "id": 359714,
  "title": "Validation - how not to OVERFIT here!",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/359714",
  "author_name": "",
  "post_date": "2022-10-13T09:50:31.094359900Z",
  "votes": 33,
  "comment_count": 13,
  "views": 0,
  "content": "<p>In the majority of competitions and data science classification problems  we are usually use several common validation techniques:</p>\n<ul>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html\" target=\"_blank\"><code>train_test_split</code></a></li>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.KFold.html\" target=\"_blank\"><code>KFold</code></a></li>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\"><code>StratifiedKFold</code></a><br>\netc.</li>\n</ul>\n<p>All of them select some part of the dataset randomly (or randomly stratified by target) for the validation purposes and it works fine if the rows in initial dataset are INDEPENDENT. But what we have in our dataset? We have rows, which are grouped using <code>game_num</code> and <code>event_id</code> values and sorted by <code>event_time</code> - so they are NOT independent. Let's take a look on the rows of the <code>train_0.csv</code> - they have the same target for the same <code>game_num</code> and <code>event_id</code> and slightly changed <code>event_time</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc4979ec50a63d3b9d2fc967943b1c683%2Ftrain_df_rows_dependence.png?generation=1665653466912486&amp;alt=media\" alt=\"\"></p>\n<p>So if you split the data randomly, the rows with the same <code>game_num</code> and <code>event_id</code> and, for example, timestamps 3 and 5 are in train data, while timestamp 4 is in test. What do you can say about the target for timestamp 4? Obviously it is almost 100% equal to the  timestamps 3 and 5 targets - <strong>this is the target leak and this leads to model overfittting.</strong></p>\n<p>To solve this problem you can use so called <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold</a> - the validation technique which know the dependencies between rows in the dataset and can split it in the right manner using the specified <code>group</code>. If you use <code>game_num</code> as group, you can eliminate the target leak and receive the real model quality. Another variant (similar to <code>train_test_split</code>) is to separate <code>game_num</code> array into 2 parts and select to train/test the rows which has <code>game_num</code> from the first/second part - the example of this technique is used by <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> and me in the kernels like <a href=\"https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta\" target=\"_blank\">this in cell 4</a>.</p>\n<p>Be careful with your validation and decrease the gap between validation and test scores - that's important to choose the right model for private LB.</p>\n<p>Alex</p>",
  "messages": [
    {
      "id": "1985402",
      "postDate": "10/13/2022 09:50:31",
      "content": "<p>In the majority of competitions and data science classification problems  we are usually use several common validation techniques:</p>\n<ul>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html\" target=\"_blank\"><code>train_test_split</code></a></li>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.KFold.html\" target=\"_blank\"><code>KFold</code></a></li>\n<li><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\"><code>StratifiedKFold</code></a><br>\netc.</li>\n</ul>\n<p>All of them select some part of the dataset randomly (or randomly stratified by target) for the validation purposes and it works fine if the rows in initial dataset are INDEPENDENT. But what we have in our dataset? We have rows, which are grouped using <code>game_num</code> and <code>event_id</code> values and sorted by <code>event_time</code> - so they are NOT independent. Let's take a look on the rows of the <code>train_0.csv</code> - they have the same target for the same <code>game_num</code> and <code>event_id</code> and slightly changed <code>event_time</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc4979ec50a63d3b9d2fc967943b1c683%2Ftrain_df_rows_dependence.png?generation=1665653466912486&amp;alt=media\" alt=\"\"></p>\n<p>So if you split the data randomly, the rows with the same <code>game_num</code> and <code>event_id</code> and, for example, timestamps 3 and 5 are in train data, while timestamp 4 is in test. What do you can say about the target for timestamp 4? Obviously it is almost 100% equal to the  timestamps 3 and 5 targets - <strong>this is the target leak and this leads to model overfittting.</strong></p>\n<p>To solve this problem you can use so called <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html\" target=\"_blank\">GroupKFold</a> - the validation technique which know the dependencies between rows in the dataset and can split it in the right manner using the specified <code>group</code>. If you use <code>game_num</code> as group, you can eliminate the target leak and receive the real model quality. Another variant (similar to <code>train_test_split</code>) is to separate <code>game_num</code> array into 2 parts and select to train/test the rows which has <code>game_num</code> from the first/second part - the example of this technique is used by <a href=\"https://www.kaggle.com/paddykb\" target=\"_blank\">@paddykb</a> and me in the kernels like <a href=\"https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta\" target=\"_blank\">this in cell 4</a>.</p>\n<p>Be careful with your validation and decrease the gap between validation and test scores - that's important to choose the right model for private LB.</p>\n<p>Alex</p>",
      "rawMarkdown": "In the majority of competitions and data science classification problems  we are usually use several common validation techniques:\n- [`train_test_split`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html)\n- [`KFold`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.KFold.html)\n- [`StratifiedKFold`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html)\netc.\n\nAll of them select some part of the dataset randomly (or randomly stratified by target) for the validation purposes and it works fine if the rows in initial dataset are INDEPENDENT. But what we have in our dataset? We have rows, which are grouped using `game_num` and `event_id` values and sorted by `event_time` - so they are NOT independent. Let's take a look on the rows of the `train_0.csv` - they have the same target for the same `game_num` and `event_id` and slightly changed `event_time`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc4979ec50a63d3b9d2fc967943b1c683%2Ftrain_df_rows_dependence.png?generation=1665653466912486&alt=media)\n\nSo if you split the data randomly, the rows with the same `game_num` and `event_id` and, for example, timestamps 3 and 5 are in train data, while timestamp 4 is in test. What do you can say about the target for timestamp 4? Obviously it is almost 100% equal to the  timestamps 3 and 5 targets - **this is the target leak and this leads to model overfittting.**\n\nTo solve this problem you can use so called [GroupKFold](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html) - the validation technique which know the dependencies between rows in the dataset and can split it in the right manner using the specified `group`. If you use `game_num` as group, you can eliminate the target leak and receive the real model quality. Another variant (similar to `train_test_split`) is to separate `game_num` array into 2 parts and select to train/test the rows which has `game_num` from the first/second part - the example of this technique is used by @paddykb and me in the kernels like [this in cell 4](https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta).\n\nBe careful with your validation and decrease the gap between validation and test scores - that's important to choose the right model for private LB.\n\nAlex",
      "votes": null
    },
    {
      "id": "1986367",
      "postDate": "10/14/2022 04:20:00",
      "content": "<p>Hi Alex great explanation! I love been in sync with other people. Just to complement, I made a <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356786\" target=\"_blank\">similar post</a> a while ago, maybe some of these ideas can be included here too.</p>",
      "rawMarkdown": "Hi Alex great explanation! I love been in sync with other people. Just to complement, I made a [similar post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356786) a while ago, maybe some of these ideas can be included here too.",
      "votes": null
    },
    {
      "id": "1986842",
      "postDate": "10/14/2022 10:38:11",
      "content": "<p><a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> yep, you are right - haven't found your post before writing this one. So we are in one boat, man :)</p>",
      "rawMarkdown": "jcaliz yep, you are right - haven't found your post before writing this one. So we are in one boat, man :)",
      "votes": null
    },
    {
      "id": "1987696",
      "postDate": "10/14/2022 19:45:12",
      "content": "<p>Oh dear… So that might be why we are seeing validation losses of around ~0.18 in most kernels but then leaderboard scores of about ~0.19-0.22!</p>",
      "rawMarkdown": "Oh dear... So that might be why we are seeing validation losses of around ~0.18 in most kernels but then leaderboard scores of about ~0.19-0.22!",
      "votes": null
    },
    {
      "id": "1987856",
      "postDate": "10/14/2022 21:15:54",
      "content": "<p>Hi Alex, you are totally right about data leakage caused by splitting randomly.</p>\n<p>These days I've been thinking about another issue that may arise from the structure of the dataset,  should we take every sample available.<br>\nWith a reasoning similar to the one you made we can say that giving to the model samples 3,4 and 5 we are giving really similar samples to the network, so we are reinforcing a signal by giving many similar samples. Is this good for the model, probably not too much.</p>\n<p>I'm looking into approaches like subsampling after grouping same game numbers, taking like you did the notebook used created by paddykb in his kernel we could reduce training data in this way:<br>\ntrain_feature_tensor = torch.tensor(<br>\n    df_train.query(\"game_num in <a href=\"https://www.kaggle.com/train\" target=\"_blank\">@train</a>_game_nums\")[[\"game_num\"]+features].groupby(\"game_num\").sample(frac=0.4, random_state=161194)[features].to_numpy())</p>\n<p>this code sample for each game_num around 40% of the steps of that game, this should allow us to get less similar samples.</p>\n<p>Another thing I was thinking during the last days is the need for randomly shuffling the data, especially with batch normalizations in a model having many similar samples grouped may lead to moving averages that keeps moving around, with shuffled data this problem should be reduced and the model performance should improve.</p>",
      "rawMarkdown": "Hi Alex, you are totally right about data leakage caused by splitting randomly.\n\nThese days I've been thinking about another issue that may arise from the structure of the dataset,  should we take every sample available.\nWith a reasoning similar to the one you made we can say that giving to the model samples 3,4 and 5 we are giving really similar samples to the network, so we are reinforcing a signal by giving many similar samples. Is this good for the model, probably not too much.\n\nI'm looking into approaches like subsampling after grouping same game numbers, taking like you did the notebook used created by paddykb in his kernel we could reduce training data in this way:\ntrain_feature_tensor = torch.tensor(\n    df_train.query(\"game_num in @train_game_nums\")[[\"game_num\"]+features].groupby(\"game_num\").sample(frac=0.4, random_state=161194)[features].to_numpy())\n\nthis code sample for each game_num around 40% of the steps of that game, this should allow us to get less similar samples.\n\nAnother thing I was thinking during the last days is the need for randomly shuffling the data, especially with batch normalizations in a model having many similar samples grouped may lead to moving averages that keeps moving around, with shuffled data this problem should be reduced and the model performance should improve.",
      "votes": null
    },
    {
      "id": "1987860",
      "postDate": "10/14/2022 21:18:14",
      "content": "<p>Yep, that's the case!! That's why me and <a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> stated it here to prevent troubles on the private LB :)</p>",
      "rawMarkdown": "Yep, that's the case!! That's why me and @jcaliz stated it here to prevent troubles on the private LB :)",
      "votes": null
    },
    {
      "id": "1987861",
      "postDate": "10/14/2022 21:22:22",
      "content": "<p>I feel like I should have noticed this non-independence situation!! Thank you !</p>",
      "rawMarkdown": "I feel like I should have noticed this non-independence situation!! Thank you !",
      "votes": null
    },
    {
      "id": "1987872",
      "postDate": "10/14/2022 21:34:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pietromaldini1\" target=\"_blank\">@pietromaldini1</a>,</p>\n<p>First of all, thanks Great comment! You deserve my upvote for that!</p>\n<p>For the first idea - in my opinion, it is not so good to lose the data using sampling because of 2 reasons:</p>\n<ul>\n<li>each object can give you necessary information as it has its own label and your model should know it </li>\n<li>during sampling you can change the mean target and the LogLoss metric can be dramatically affected by that (so you can receive worse scores at the end)<br>\nMaybe (I haven't checked that, it is just my idea) you can <strong>use the weights for the objects</strong> to weight more each 5th-10th (this period can be optimized on validation) object to make your dataset more diverse during training.</li>\n</ul>\n<p>For the second one - I think that the data is already shuffled well using different techniques:</p>\n<ul>\n<li>The rows for the batch are coming from different <code>game_num</code>s and even different <code>event_id</code>s</li>\n<li>We are using random augmentations like players shuffling, flips by X and Y axes so we make our dataset more diverse</li>\n</ul>\n<p>Hope this can help you with your thoughts.</p>\n<p>Alex</p>",
      "rawMarkdown": "Hi @pietromaldini1,\n\nFirst of all, thanks Great comment! You deserve my upvote for that!\n\nFor the first idea - in my opinion, it is not so good to lose the data using sampling because of 2 reasons:\n- each object can give you necessary information as it has its own label and your model should know it \n- during sampling you can change the mean target and the LogLoss metric can be dramatically affected by that (so you can receive worse scores at the end)\nMaybe (I haven't checked that, it is just my idea) you can **use the weights for the objects** to weight more each 5th-10th (this period can be optimized on validation) object to make your dataset more diverse during training.\n\nFor the second one - I think that the data is already shuffled well using different techniques:\n- The rows for the batch are coming from different `game_num`s and even different `event_id`s\n- We are using random augmentations like players shuffling, flips by X and Y axes so we make our dataset more diverse\n\nHope this can help you with your thoughts.\n\nAlex",
      "votes": null
    },
    {
      "id": "1987896",
      "postDate": "10/14/2022 22:00:54",
      "content": "<p>Thank you for your quick and very long answer.<br>\nI'll definitely consider your comment in my next attempts to this challenge. </p>\n<p>For now I'll wait my experiments about data sampling to see how it changes the performance of the model, it probably is a dead end but analyzing how the model learn with less data may give some insights of which data points are important.</p>\n<p>Random sampling is highly likely to fail and reduce performance, but a dataset distillation approach may prove to be beneficial to the task and may allow the use of a greater number of features in training phase.</p>\n<p>I'll post an update on the discussions if I find anything interesting.</p>",
      "rawMarkdown": "Thank you for your quick and very long answer.\nI'll definitely consider your comment in my next attempts to this challenge. \n\nFor now I'll wait my experiments about data sampling to see how it changes the performance of the model, it probably is a dead end but analyzing how the model learn with less data may give some insights of which data points are important.\n\nRandom sampling is highly likely to fail and reduce performance, but a dataset distillation approach may prove to be beneficial to the task and may allow the use of a greater number of features in training phase.\n\nI'll post an update on the discussions if I find anything interesting.",
      "votes": null
    },
    {
      "id": "1989186",
      "postDate": "10/15/2022 19:02:10",
      "content": "<p>Hi Pietro, I'm wondering the same thing about your first idea. I've been working on this competition for a couple days and after some exploration I want to make a sample and work with it to save time in my future experiments. My main idea right now is to split events in time intervals of, say, three seconds and sample one or two rows from each interval. My assumption is that the game would be different enough after a couple of seconds to create a more diverse sample. Also, I want to make a code where I parametrize the number of games, the number of events and the interval size that I sample. Then, my objective is to figure out the best parameters using a simple model without fine-tuning.</p>\n<p>I don't know if it's going to pay off (probably not haha) but I think it's a cool experiment and it relates to real life, where you don't always have enough time or computing power to use large datasets or DNNs.  </p>",
      "rawMarkdown": "Hi Pietro, I'm wondering the same thing about your first idea. I've been working on this competition for a couple days and after some exploration I want to make a sample and work with it to save time in my future experiments. My main idea right now is to split events in time intervals of, say, three seconds and sample one or two rows from each interval. My assumption is that the game would be different enough after a couple of seconds to create a more diverse sample. Also, I want to make a code where I parametrize the number of games, the number of events and the interval size that I sample. Then, my objective is to figure out the best parameters using a simple model without fine-tuning.\n\nI don't know if it's going to pay off (probably not haha) but I think it's a cool experiment and it relates to real life, where you don't always have enough time or computing power to use large datasets or DNNs.",
      "votes": null
    },
    {
      "id": "1989633",
      "postDate": "10/16/2022 05:44:11",
      "content": "<p>osm nice work sir </p>",
      "rawMarkdown": "osm nice work sir",
      "votes": null
    },
    {
      "id": "1991060",
      "postDate": "10/17/2022 00:28:04",
      "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">@alexryzhkov</a> !</p>",
      "rawMarkdown": "Thank you for sharing @alexryzhkov !",
      "votes": null
    },
    {
      "id": "2001931",
      "postDate": "10/24/2022 12:08:25",
      "content": "<p>Great Information Alex,<br>\nI am wondering isn't <code>event_id</code> enough for grouping and would be more flexible to work with</p>",
      "rawMarkdown": "Great Information Alex,\nI am wondering isn't `event_id` enough for grouping and would be more flexible to work with",
      "votes": null
    },
    {
      "id": "2001988",
      "postDate": "10/24/2022 12:55:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mohamedmagdy11\" target=\"_blank\">@mohamedmagdy11</a>,</p>\n<p>I think that <code>event_id</code> is almost enough - in case of using it the only leak is the same teams in different folds. But as you have another <code>game_num</code>s in test set (I hope so) it will be more safe to use <code>game_num</code> as group</p>",
      "rawMarkdown": "Hi @mohamedmagdy11,\n\nI think that `event_id` is almost enough - in case of using it the only leak is the same teams in different folds. But as you have another `game_num`s in test set (I hope so) it will be more safe to use `game_num` as group",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1986367,
      "author_name": "jcaliz",
      "author_url": "",
      "post_date": "10/14/2022 04:20:00",
      "content": "<p>Hi Alex great explanation! I love been in sync with other people. Just to complement, I made a <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356786\" target=\"_blank\">similar post</a> a while ago, maybe some of these ideas can be included here too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1986842,
          "author_name": "alexryzhkov",
          "author_url": "",
          "post_date": "10/14/2022 10:38:11",
          "content": "<p><a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> yep, you are right - haven't found your post before writing this one. So we are in one boat, man :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1987696,
      "author_name": "fabianbong",
      "author_url": "",
      "post_date": "10/14/2022 19:45:12",
      "content": "<p>Oh dear… So that might be why we are seeing validation losses of around ~0.18 in most kernels but then leaderboard scores of about ~0.19-0.22!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1987860,
          "author_name": "alexryzhkov",
          "author_url": "",
          "post_date": "10/14/2022 21:18:14",
          "content": "<p>Yep, that's the case!! That's why me and <a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> stated it here to prevent troubles on the private LB :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1987861,
          "author_name": "fabianbong",
          "author_url": "",
          "post_date": "10/14/2022 21:22:22",
          "content": "<p>I feel like I should have noticed this non-independence situation!! Thank you !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1987856,
      "author_name": "pietromaldini1",
      "author_url": "",
      "post_date": "10/14/2022 21:15:54",
      "content": "<p>Hi Alex, you are totally right about data leakage caused by splitting randomly.</p>\n<p>These days I've been thinking about another issue that may arise from the structure of the dataset,  should we take every sample available.<br>\nWith a reasoning similar to the one you made we can say that giving to the model samples 3,4 and 5 we are giving really similar samples to the network, so we are reinforcing a signal by giving many similar samples. Is this good for the model, probably not too much.</p>\n<p>I'm looking into approaches like subsampling after grouping same game numbers, taking like you did the notebook used created by paddykb in his kernel we could reduce training data in this way:<br>\ntrain_feature_tensor = torch.tensor(<br>\n    df_train.query(\"game_num in <a href=\"https://www.kaggle.com/train\" target=\"_blank\">@train</a>_game_nums\")[[\"game_num\"]+features].groupby(\"game_num\").sample(frac=0.4, random_state=161194)[features].to_numpy())</p>\n<p>this code sample for each game_num around 40% of the steps of that game, this should allow us to get less similar samples.</p>\n<p>Another thing I was thinking during the last days is the need for randomly shuffling the data, especially with batch normalizations in a model having many similar samples grouped may lead to moving averages that keeps moving around, with shuffled data this problem should be reduced and the model performance should improve.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1987872,
          "author_name": "alexryzhkov",
          "author_url": "",
          "post_date": "10/14/2022 21:34:20",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pietromaldini1\" target=\"_blank\">@pietromaldini1</a>,</p>\n<p>First of all, thanks Great comment! You deserve my upvote for that!</p>\n<p>For the first idea - in my opinion, it is not so good to lose the data using sampling because of 2 reasons:</p>\n<ul>\n<li>each object can give you necessary information as it has its own label and your model should know it </li>\n<li>during sampling you can change the mean target and the LogLoss metric can be dramatically affected by that (so you can receive worse scores at the end)<br>\nMaybe (I haven't checked that, it is just my idea) you can <strong>use the weights for the objects</strong> to weight more each 5th-10th (this period can be optimized on validation) object to make your dataset more diverse during training.</li>\n</ul>\n<p>For the second one - I think that the data is already shuffled well using different techniques:</p>\n<ul>\n<li>The rows for the batch are coming from different <code>game_num</code>s and even different <code>event_id</code>s</li>\n<li>We are using random augmentations like players shuffling, flips by X and Y axes so we make our dataset more diverse</li>\n</ul>\n<p>Hope this can help you with your thoughts.</p>\n<p>Alex</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1987896,
          "author_name": "pietromaldini1",
          "author_url": "",
          "post_date": "10/14/2022 22:00:54",
          "content": "<p>Thank you for your quick and very long answer.<br>\nI'll definitely consider your comment in my next attempts to this challenge. </p>\n<p>For now I'll wait my experiments about data sampling to see how it changes the performance of the model, it probably is a dead end but analyzing how the model learn with less data may give some insights of which data points are important.</p>\n<p>Random sampling is highly likely to fail and reduce performance, but a dataset distillation approach may prove to be beneficial to the task and may allow the use of a greater number of features in training phase.</p>\n<p>I'll post an update on the discussions if I find anything interesting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1989186,
          "author_name": "mateuscco",
          "author_url": "",
          "post_date": "10/15/2022 19:02:10",
          "content": "<p>Hi Pietro, I'm wondering the same thing about your first idea. I've been working on this competition for a couple days and after some exploration I want to make a sample and work with it to save time in my future experiments. My main idea right now is to split events in time intervals of, say, three seconds and sample one or two rows from each interval. My assumption is that the game would be different enough after a couple of seconds to create a more diverse sample. Also, I want to make a code where I parametrize the number of games, the number of events and the interval size that I sample. Then, my objective is to figure out the best parameters using a simple model without fine-tuning.</p>\n<p>I don't know if it's going to pay off (probably not haha) but I think it's a cool experiment and it relates to real life, where you don't always have enough time or computing power to use large datasets or DNNs.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1989633,
      "author_name": "anoshkarokhar",
      "author_url": "",
      "post_date": "10/16/2022 05:44:11",
      "content": "<p>osm nice work sir </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1991060,
      "author_name": "arti1117",
      "author_url": "",
      "post_date": "10/17/2022 00:28:04",
      "content": "<p>Thank you for sharing <a href=\"https://www.kaggle.com/alexryzhkov\" target=\"_blank\">@alexryzhkov</a> !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2001931,
      "author_name": "mohamedmagdy11",
      "author_url": "",
      "post_date": "10/24/2022 12:08:25",
      "content": "<p>Great Information Alex,<br>\nI am wondering isn't <code>event_id</code> enough for grouping and would be more flexible to work with</p>",
      "votes": null,
      "replies": [
        {
          "id": 2001988,
          "author_name": "alexryzhkov",
          "author_url": "",
          "post_date": "10/24/2022 12:55:06",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mohamedmagdy11\" target=\"_blank\">@mohamedmagdy11</a>,</p>\n<p>I think that <code>event_id</code> is almost enough - in case of using it the only leak is the same teams in different folds. But as you have another <code>game_num</code>s in test set (I hope so) it will be more safe to use <code>game_num</code> as group</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1985402": "In the majority of competitions and data science classification problems  we are usually use several common validation techniques:\n- [`train_test_split`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html)\n- [`KFold`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.KFold.html)\n- [`StratifiedKFold`](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html)\netc.\n\nAll of them select some part of the dataset randomly (or randomly stratified by target) for the validation purposes and it works fine if the rows in initial dataset are INDEPENDENT. But what we have in our dataset? We have rows, which are grouped using `game_num` and `event_id` values and sorted by `event_time` - so they are NOT independent. Let's take a look on the rows of the `train_0.csv` - they have the same target for the same `game_num` and `event_id` and slightly changed `event_time`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19099%2Fc4979ec50a63d3b9d2fc967943b1c683%2Ftrain_df_rows_dependence.png?generation=1665653466912486&alt=media)\n\nSo if you split the data randomly, the rows with the same `game_num` and `event_id` and, for example, timestamps 3 and 5 are in train data, while timestamp 4 is in test. What do you can say about the target for timestamp 4? Obviously it is almost 100% equal to the  timestamps 3 and 5 targets - **this is the target leak and this leads to model overfittting.**\n\nTo solve this problem you can use so called [GroupKFold](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html) - the validation technique which know the dependencies between rows in the dataset and can split it in the right manner using the specified `group`. If you use `game_num` as group, you can eliminate the target leak and receive the real model quality. Another variant (similar to `train_test_split`) is to separate `game_num` array into 2 parts and select to train/test the rows which has `game_num` from the first/second part - the example of this technique is used by @paddykb and me in the kernels like [this in cell 4](https://www.kaggle.com/code/alexryzhkov/tps-2022-10-fastai-with-multistart-and-tta).\n\nBe careful with your validation and decrease the gap between validation and test scores - that's important to choose the right model for private LB.\n\nAlex",
    "1986367": "Hi Alex great explanation! I love been in sync with other people. Just to complement, I made a [similar post](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356786) a while ago, maybe some of these ideas can be included here too.",
    "1986842": "jcaliz yep, you are right - haven't found your post before writing this one. So we are in one boat, man :)",
    "1987696": "Oh dear... So that might be why we are seeing validation losses of around ~0.18 in most kernels but then leaderboard scores of about ~0.19-0.22!",
    "1987856": "Hi Alex, you are totally right about data leakage caused by splitting randomly.\n\nThese days I've been thinking about another issue that may arise from the structure of the dataset,  should we take every sample available.\nWith a reasoning similar to the one you made we can say that giving to the model samples 3,4 and 5 we are giving really similar samples to the network, so we are reinforcing a signal by giving many similar samples. Is this good for the model, probably not too much.\n\nI'm looking into approaches like subsampling after grouping same game numbers, taking like you did the notebook used created by paddykb in his kernel we could reduce training data in this way:\ntrain_feature_tensor = torch.tensor(\n    df_train.query(\"game_num in @train_game_nums\")[[\"game_num\"]+features].groupby(\"game_num\").sample(frac=0.4, random_state=161194)[features].to_numpy())\n\nthis code sample for each game_num around 40% of the steps of that game, this should allow us to get less similar samples.\n\nAnother thing I was thinking during the last days is the need for randomly shuffling the data, especially with batch normalizations in a model having many similar samples grouped may lead to moving averages that keeps moving around, with shuffled data this problem should be reduced and the model performance should improve.",
    "1987860": "Yep, that's the case!! That's why me and @jcaliz stated it here to prevent troubles on the private LB :)",
    "1987861": "I feel like I should have noticed this non-independence situation!! Thank you !",
    "1987872": "Hi @pietromaldini1,\n\nFirst of all, thanks Great comment! You deserve my upvote for that!\n\nFor the first idea - in my opinion, it is not so good to lose the data using sampling because of 2 reasons:\n- each object can give you necessary information as it has its own label and your model should know it \n- during sampling you can change the mean target and the LogLoss metric can be dramatically affected by that (so you can receive worse scores at the end)\nMaybe (I haven't checked that, it is just my idea) you can **use the weights for the objects** to weight more each 5th-10th (this period can be optimized on validation) object to make your dataset more diverse during training.\n\nFor the second one - I think that the data is already shuffled well using different techniques:\n- The rows for the batch are coming from different `game_num`s and even different `event_id`s\n- We are using random augmentations like players shuffling, flips by X and Y axes so we make our dataset more diverse\n\nHope this can help you with your thoughts.\n\nAlex",
    "1987896": "Thank you for your quick and very long answer.\nI'll definitely consider your comment in my next attempts to this challenge. \n\nFor now I'll wait my experiments about data sampling to see how it changes the performance of the model, it probably is a dead end but analyzing how the model learn with less data may give some insights of which data points are important.\n\nRandom sampling is highly likely to fail and reduce performance, but a dataset distillation approach may prove to be beneficial to the task and may allow the use of a greater number of features in training phase.\n\nI'll post an update on the discussions if I find anything interesting.",
    "1989186": "Hi Pietro, I'm wondering the same thing about your first idea. I've been working on this competition for a couple days and after some exploration I want to make a sample and work with it to save time in my future experiments. My main idea right now is to split events in time intervals of, say, three seconds and sample one or two rows from each interval. My assumption is that the game would be different enough after a couple of seconds to create a more diverse sample. Also, I want to make a code where I parametrize the number of games, the number of events and the interval size that I sample. Then, my objective is to figure out the best parameters using a simple model without fine-tuning.\n\nI don't know if it's going to pay off (probably not haha) but I think it's a cool experiment and it relates to real life, where you don't always have enough time or computing power to use large datasets or DNNs.",
    "1989633": "osm nice work sir",
    "1991060": "Thank you for sharing @alexryzhkov !",
    "2001931": "Great Information Alex,\nI am wondering isn't `event_id` enough for grouping and would be more flexible to work with",
    "2001988": "Hi @mohamedmagdy11,\n\nI think that `event_id` is almost enough - in case of using it the only leak is the same teams in different folds. But as you have another `game_num`s in test set (I hope so) it will be more safe to use `game_num` as group"
  },
  "source": "meta"
}