{
  "id": 420202,
  "title": "973 place solution :)",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420202",
  "author_name": "",
  "post_date": "2023-06-29T16:04:02.998204900Z",
  "votes": 18,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was mainly competing for the effeciency prize, so I went for an Lstm model on raw data with no preprocessing or manual feature engineering.<br>\nI have 2 lstms inside the same model, one that deals only with user's interactions (object_click, notification, notebook, ..) and one that deals only with conversations (cutscenes and person_click). As features I used the embeddings on the categorical features and the difference of elapsed time between 2 user's actions.</p>\n<p>My first attempt got me a 0.698 on a 5 folds CV and 0.701 on LB. With some optimizations I was able to reach 0.7 CV and 0.703 on LB.</p>\n<p>Unfortunately, all my submissions on private LB are below 0.69 and end up dropping by 905 places. I have no idea why!!! </p>",
  "messages": [
    {
      "id": "2323000",
      "postDate": "06/29/2023 16:04:02",
      "content": "<p>I was mainly competing for the effeciency prize, so I went for an Lstm model on raw data with no preprocessing or manual feature engineering.<br>\nI have 2 lstms inside the same model, one that deals only with user's interactions (object_click, notification, notebook, ..) and one that deals only with conversations (cutscenes and person_click). As features I used the embeddings on the categorical features and the difference of elapsed time between 2 user's actions.</p>\n<p>My first attempt got me a 0.698 on a 5 folds CV and 0.701 on LB. With some optimizations I was able to reach 0.7 CV and 0.703 on LB.</p>\n<p>Unfortunately, all my submissions on private LB are below 0.69 and end up dropping by 905 places. I have no idea why!!! </p>",
      "rawMarkdown": "I was mainly competing for the effeciency prize, so I went for an Lstm model on raw data with no preprocessing or manual feature engineering.\nI have 2 lstms inside the same model, one that deals only with user's interactions (object_click, notification, notebook, ..) and one that deals only with conversations (cutscenes and person_click). As features I used the embeddings on the categorical features and the difference of elapsed time between 2 user's actions.\n\nMy first attempt got me a 0.698 on a 5 folds CV and 0.701 on LB. With some optimizations I was able to reach 0.7 CV and 0.703 on LB.\n\nUnfortunately, all my submissions on private LB are below 0.69 and end up dropping by 905 places. I have no idea why!!!",
      "votes": null
    },
    {
      "id": "2323014",
      "postDate": "06/29/2023 16:16:17",
      "content": "<p>Wow, this seems like the biggest shake-down I have seen from medal-zone in this competition.</p>",
      "rawMarkdown": "Wow, this seems like the biggest shake-down I have seen from medal-zone in this competition.",
      "votes": null
    },
    {
      "id": "2323337",
      "postDate": "06/29/2023 21:54:39",
      "content": "<p>there is definitely some unsolved mysteries here. I think we survived thanks to blending very different models. But our efficiency sub is a single model, and its private score is also way less than the public score,</p>",
      "rawMarkdown": "there is definitely some unsolved mysteries here. I think we survived thanks to blending very different models. But our efficiency sub is a single model, and its private score is also way less than the public score,",
      "votes": null
    },
    {
      "id": "2325863",
      "postDate": "07/01/2023 16:34:17",
      "content": "<p>Quite similar. I have one tree-based model and one LSTM-like NN model which scored 0.705 publicly, but drop 1899 places with only 0.586 privatly. The NN is the main reason. I guess some private data's sequence is different with the public one.  </p>",
      "rawMarkdown": "Quite similar. I have one tree-based model and one LSTM-like NN model which scored 0.705 publicly, but drop 1899 places with only 0.586 privatly. The NN is the main reason. I guess some private data's sequence is different with the public one.",
      "votes": null
    },
    {
      "id": "2325970",
      "postDate": "07/01/2023 18:09:54",
      "content": "<p>I guess so, I tried one lstm model that uses the difference of elapsed_time from end to start for each event_name, fqid, text_fqid values, it was not as good as using raw data on CV and public LB but it scores 0.7 on private LB.</p>\n<p>I would like to hear from the organizers if the private LB data is different from the training and public LB data</p>",
      "rawMarkdown": "I guess so, I tried one lstm model that uses the difference of elapsed_time from end to start for each event_name, fqid, text_fqid values, it was not as good as using raw data on CV and public LB but it scores 0.7 on private LB.\n\nI would like to hear from the organizers if the private LB data is different from the training and public LB data",
      "votes": null
    },
    {
      "id": "2325992",
      "postDate": "07/01/2023 18:47:02",
      "content": "<p>Could be some real mysteries due to real reasons. (How many people used the \"year\" feature? How did the private LB handle someone that played in 2021 playing again in 2023? Playing twice in Jun 2023?)</p>\n<p>But don't underestimate pure variance on small sample size due to data leak! My efficiency sub happened to get lucky with private LB increase, but could've gone the other way I'm sure</p>",
      "rawMarkdown": "Could be some real mysteries due to real reasons. (How many people used the \"year\" feature? How did the private LB handle someone that played in 2021 playing again in 2023? Playing twice in Jun 2023?)\n\nBut don't underestimate pure variance on small sample size due to data leak! My efficiency sub happened to get lucky with private LB increase, but could've gone the other way I'm sure",
      "votes": null
    },
    {
      "id": "2325994",
      "postDate": "07/01/2023 18:48:52",
      "content": "<p>Maybe a sorting issue? Did you sort your train and test data before feature engineering?</p>",
      "rawMarkdown": "Maybe a sorting issue? Did you sort your train and test data before feature engineering?",
      "votes": null
    },
    {
      "id": "2325995",
      "postDate": "07/01/2023 18:53:06",
      "content": "<p>Oh, I see no preprocessing or feature engineering in description. I do suspect it could be a sorting issue. And for some reason impact private but not public LB, which is pretty awful :/</p>",
      "rawMarkdown": "Oh, I see no preprocessing or feature engineering in description. I do suspect it could be a sorting issue. And for some reason impact private but not public LB, which is pretty awful :/",
      "votes": null
    },
    {
      "id": "2325996",
      "postDate": "07/01/2023 18:56:17",
      "content": "<p>of course I did, if you don't sort your data, you get an awful public score too</p>",
      "rawMarkdown": "of course I did, if you don't sort your data, you get an awful public score too",
      "votes": null
    },
    {
      "id": "2326742",
      "postDate": "07/02/2023 11:16:30",
      "content": "<p>It looks like that NNs fare worse on private LB than GBMs from what I see on our submissions and what I read on the forum. That's the mystery I speak about.</p>\n<p>I agree with you that variability is higher with a single model than for an ensemble, and this maybe part of the explanation too.</p>",
      "rawMarkdown": "It looks like that NNs fare worse on private LB than GBMs from what I see on our submissions and what I read on the forum. That's the mystery I speak about.\n\nI agree with you that variability is higher with a single model than for an ensemble, and this maybe part of the explanation too.",
      "votes": null
    },
    {
      "id": "2327232",
      "postDate": "07/02/2023 18:39:13",
      "content": "<p>Good point, thanks for clarifying!</p>\n<p>One thing I'll note: there are four different text version options for the game. Snark vs no snark, and another binary I forget. (I haven't noticed any write ups that spent time talking about that.) That by itself might be more likely to impact NN by construction than the GBM typical features(?) Plus maybe private was a different distribution for some reason, or even just one or two of those options, or they stopped doing all four, so all private was just one?</p>\n<p>But yeah, mystery. That's just idle speculation trying to find simplest possible answer. </p>",
      "rawMarkdown": "Good point, thanks for clarifying!\n\nOne thing I'll note: there are four different text version options for the game. Snark vs no snark, and another binary I forget. (I haven't noticed any write ups that spent time talking about that.) That by itself might be more likely to impact NN by construction than the GBM typical features(?) Plus maybe private was a different distribution for some reason, or even just one or two of those options, or they stopped doing all four, so all private was just one?\n\nBut yeah, mystery. That's just idle speculation trying to find simplest possible answer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2323014,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/29/2023 16:16:17",
      "content": "<p>Wow, this seems like the biggest shake-down I have seen from medal-zone in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323337,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/29/2023 21:54:39",
      "content": "<p>there is definitely some unsolved mysteries here. I think we survived thanks to blending very different models. But our efficiency sub is a single model, and its private score is also way less than the public score,</p>",
      "votes": null,
      "replies": [
        {
          "id": 2325992,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/01/2023 18:47:02",
          "content": "<p>Could be some real mysteries due to real reasons. (How many people used the \"year\" feature? How did the private LB handle someone that played in 2021 playing again in 2023? Playing twice in Jun 2023?)</p>\n<p>But don't underestimate pure variance on small sample size due to data leak! My efficiency sub happened to get lucky with private LB increase, but could've gone the other way I'm sure</p>",
          "votes": null,
          "replies": [
            {
              "id": 2326742,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "07/02/2023 11:16:30",
              "content": "<p>It looks like that NNs fare worse on private LB than GBMs from what I see on our submissions and what I read on the forum. That's the mystery I speak about.</p>\n<p>I agree with you that variability is higher with a single model than for an ensemble, and this maybe part of the explanation too.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2327232,
                  "author_name": "roberthatch",
                  "author_url": "",
                  "post_date": "07/02/2023 18:39:13",
                  "content": "<p>Good point, thanks for clarifying!</p>\n<p>One thing I'll note: there are four different text version options for the game. Snark vs no snark, and another binary I forget. (I haven't noticed any write ups that spent time talking about that.) That by itself might be more likely to impact NN by construction than the GBM typical features(?) Plus maybe private was a different distribution for some reason, or even just one or two of those options, or they stopped doing all four, so all private was just one?</p>\n<p>But yeah, mystery. That's just idle speculation trying to find simplest possible answer. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2325863,
      "author_name": "littlstar123",
      "author_url": "",
      "post_date": "07/01/2023 16:34:17",
      "content": "<p>Quite similar. I have one tree-based model and one LSTM-like NN model which scored 0.705 publicly, but drop 1899 places with only 0.586 privatly. The NN is the main reason. I guess some private data's sequence is different with the public one.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2325970,
          "author_name": "mchahhou",
          "author_url": "",
          "post_date": "07/01/2023 18:09:54",
          "content": "<p>I guess so, I tried one lstm model that uses the difference of elapsed_time from end to start for each event_name, fqid, text_fqid values, it was not as good as using raw data on CV and public LB but it scores 0.7 on private LB.</p>\n<p>I would like to hear from the organizers if the private LB data is different from the training and public LB data</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2325994,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/01/2023 18:48:52",
      "content": "<p>Maybe a sorting issue? Did you sort your train and test data before feature engineering?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2325995,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "07/01/2023 18:53:06",
          "content": "<p>Oh, I see no preprocessing or feature engineering in description. I do suspect it could be a sorting issue. And for some reason impact private but not public LB, which is pretty awful :/</p>",
          "votes": null,
          "replies": [
            {
              "id": 2325996,
              "author_name": "mchahhou",
              "author_url": "",
              "post_date": "07/01/2023 18:56:17",
              "content": "<p>of course I did, if you don't sort your data, you get an awful public score too</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2323000": "I was mainly competing for the effeciency prize, so I went for an Lstm model on raw data with no preprocessing or manual feature engineering.\nI have 2 lstms inside the same model, one that deals only with user's interactions (object_click, notification, notebook, ..) and one that deals only with conversations (cutscenes and person_click). As features I used the embeddings on the categorical features and the difference of elapsed time between 2 user's actions.\n\nMy first attempt got me a 0.698 on a 5 folds CV and 0.701 on LB. With some optimizations I was able to reach 0.7 CV and 0.703 on LB.\n\nUnfortunately, all my submissions on private LB are below 0.69 and end up dropping by 905 places. I have no idea why!!!",
    "2323014": "Wow, this seems like the biggest shake-down I have seen from medal-zone in this competition.",
    "2323337": "there is definitely some unsolved mysteries here. I think we survived thanks to blending very different models. But our efficiency sub is a single model, and its private score is also way less than the public score,",
    "2325863": "Quite similar. I have one tree-based model and one LSTM-like NN model which scored 0.705 publicly, but drop 1899 places with only 0.586 privatly. The NN is the main reason. I guess some private data's sequence is different with the public one.",
    "2325970": "I guess so, I tried one lstm model that uses the difference of elapsed_time from end to start for each event_name, fqid, text_fqid values, it was not as good as using raw data on CV and public LB but it scores 0.7 on private LB.\n\nI would like to hear from the organizers if the private LB data is different from the training and public LB data",
    "2325992": "Could be some real mysteries due to real reasons. (How many people used the \"year\" feature? How did the private LB handle someone that played in 2021 playing again in 2023? Playing twice in Jun 2023?)\n\nBut don't underestimate pure variance on small sample size due to data leak! My efficiency sub happened to get lucky with private LB increase, but could've gone the other way I'm sure",
    "2325994": "Maybe a sorting issue? Did you sort your train and test data before feature engineering?",
    "2325995": "Oh, I see no preprocessing or feature engineering in description. I do suspect it could be a sorting issue. And for some reason impact private but not public LB, which is pretty awful :/",
    "2325996": "of course I did, if you don't sort your data, you get an awful public score too",
    "2326742": "It looks like that NNs fare worse on private LB than GBMs from what I see on our submissions and what I read on the forum. That's the mystery I speak about.\n\nI agree with you that variability is higher with a single model than for an ensemble, and this maybe part of the explanation too.",
    "2327232": "Good point, thanks for clarifying!\n\nOne thing I'll note: there are four different text version options for the game. Snark vs no snark, and another binary I forget. (I haven't noticed any write ups that spent time talking about that.) That by itself might be more likely to impact NN by construction than the GBM typical features(?) Plus maybe private was a different distribution for some reason, or even just one or two of those options, or they stopped doing all four, so all private was just one?\n\nBut yeah, mystery. That's just idle speculation trying to find simplest possible answer."
  },
  "source": "meta"
}