{
  "id": 413002,
  "title": "LB score is very low with new API",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/413002",
  "author_name": "",
  "post_date": "2023-05-26T09:51:05.150045900Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>As has been pointed out in this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">discussion</a>, the new API for submission has a number of problems.<br>\nThere are reports that sorting the index and question columns has made the scores correct, but my LB scores remain low. (CV=0.688, VB=0.612)<br>\nI am struggling to find the cause but am having a very hard time without any clues.<br>\nHas anyone else experienced a similar phenomenon?<br>\nFor your reference, I will share some puzzling points found in some experiments.</p>\n<p>・Score corruption occurs only in models using aggregate features, not in NN models that use log data as is.<br>\n・The results do not change depending on whether the index is sorted or not.<br>\n・In the same pipeline, some score destruction occurs depending on the number of features, while others do not. (All worked before).<br>\n・After the new API was released, the same phenomenon started to occur with the old API.<br>\n・Submission error occurs when executing the following code that does sorting of index.</p>\n<pre><code>df = df.groupby(['session_id']).apply(lambda df: df.sort_values('index'))\ndf = df.reset_index(drop=True)\n</code></pre>",
  "messages": [
    {
      "id": "2274854",
      "postDate": "05/26/2023 09:51:05",
      "content": "<p>As has been pointed out in this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">discussion</a>, the new API for submission has a number of problems.<br>\nThere are reports that sorting the index and question columns has made the scores correct, but my LB scores remain low. (CV=0.688, VB=0.612)<br>\nI am struggling to find the cause but am having a very hard time without any clues.<br>\nHas anyone else experienced a similar phenomenon?<br>\nFor your reference, I will share some puzzling points found in some experiments.</p>\n<p>・Score corruption occurs only in models using aggregate features, not in NN models that use log data as is.<br>\n・The results do not change depending on whether the index is sorted or not.<br>\n・In the same pipeline, some score destruction occurs depending on the number of features, while others do not. (All worked before).<br>\n・After the new API was released, the same phenomenon started to occur with the old API.<br>\n・Submission error occurs when executing the following code that does sorting of index.</p>\n<pre><code>df = df.groupby(['session_id']).apply(lambda df: df.sort_values('index'))\ndf = df.reset_index(drop=True)\n</code></pre>",
      "rawMarkdown": "As has been pointed out in this [discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512), the new API for submission has a number of problems.\nThere are reports that sorting the index and question columns has made the scores correct, but my LB scores remain low. (CV=0.688, VB=0.612)\nI am struggling to find the cause but am having a very hard time without any clues.\nHas anyone else experienced a similar phenomenon?\nFor your reference, I will share some puzzling points found in some experiments.\n\n・Score corruption occurs only in models using aggregate features, not in NN models that use log data as is.\n・The results do not change depending on whether the index is sorted or not.\n・In the same pipeline, some score destruction occurs depending on the number of features, while others do not. (All worked before).\n・After the new API was released, the same phenomenon started to occur with the old API.\n・Submission error occurs when executing the following code that does sorting of index.\n```\ndf = df.groupby(['session_id']).apply(lambda df: df.sort_values('index'))\ndf = df.reset_index(drop=True)\n```",
      "votes": null
    },
    {
      "id": "2276894",
      "postDate": "05/27/2023 09:41:01",
      "content": "<p>Just use old api :) <br>\nThe probability that Phil will react soon is low</p>",
      "rawMarkdown": "Just use old api :) \nThe probability that Phil will react soon is low",
      "votes": null
    },
    {
      "id": "2280750",
      "postDate": "05/30/2023 10:29:22",
      "content": "<p>I have solved this problem!<br>\nIn fact, my final problem was not caused by the API, but by the fact that the feature calculation was not kept reproducible.<br>\nIn other words, the sorting process mentioned in this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">discussion</a> was sufficient to keep the LB scores accurate with the new API.<br>\nFor those who are similarly stuck on this issue, a reproducible script (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413004\" target=\"_blank\">here</a>) is also available for your reference.<br>\nSubmission Error always gives me nightmares…<br>\nI hope that as many people as possible will be able to solve their problems.</p>",
      "rawMarkdown": "I have solved this problem!\nIn fact, my final problem was not caused by the API, but by the fact that the feature calculation was not kept reproducible.\nIn other words, the sorting process mentioned in this [discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512) was sufficient to keep the LB scores accurate with the new API.\nFor those who are similarly stuck on this issue, a reproducible script ([here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413004)) is also available for your reference.\nSubmission Error always gives me nightmares...\nI hope that as many people as possible will be able to solve their problems.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2276894,
      "author_name": "kvlmll",
      "author_url": "",
      "post_date": "05/27/2023 09:41:01",
      "content": "<p>Just use old api :) <br>\nThe probability that Phil will react soon is low</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2280750,
      "author_name": "jinmiyashita",
      "author_url": "",
      "post_date": "05/30/2023 10:29:22",
      "content": "<p>I have solved this problem!<br>\nIn fact, my final problem was not caused by the API, but by the fact that the feature calculation was not kept reproducible.<br>\nIn other words, the sorting process mentioned in this <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">discussion</a> was sufficient to keep the LB scores accurate with the new API.<br>\nFor those who are similarly stuck on this issue, a reproducible script (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413004\" target=\"_blank\">here</a>) is also available for your reference.<br>\nSubmission Error always gives me nightmares…<br>\nI hope that as many people as possible will be able to solve their problems.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2274854": "As has been pointed out in this [discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512), the new API for submission has a number of problems.\nThere are reports that sorting the index and question columns has made the scores correct, but my LB scores remain low. (CV=0.688, VB=0.612)\nI am struggling to find the cause but am having a very hard time without any clues.\nHas anyone else experienced a similar phenomenon?\nFor your reference, I will share some puzzling points found in some experiments.\n\n・Score corruption occurs only in models using aggregate features, not in NN models that use log data as is.\n・The results do not change depending on whether the index is sorted or not.\n・In the same pipeline, some score destruction occurs depending on the number of features, while others do not. (All worked before).\n・After the new API was released, the same phenomenon started to occur with the old API.\n・Submission error occurs when executing the following code that does sorting of index.\n```\ndf = df.groupby(['session_id']).apply(lambda df: df.sort_values('index'))\ndf = df.reset_index(drop=True)\n```",
    "2276894": "Just use old api :) \nThe probability that Phil will react soon is low",
    "2280750": "I have solved this problem!\nIn fact, my final problem was not caused by the API, but by the fact that the feature calculation was not kept reproducible.\nIn other words, the sorting process mentioned in this [discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512) was sufficient to keep the LB scores accurate with the new API.\nFor those who are similarly stuck on this issue, a reproducible script ([here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413004)) is also available for your reference.\nSubmission Error always gives me nightmares...\nI hope that as many people as possible will be able to solve their problems."
  },
  "source": "meta"
}