{
  "id": 415215,
  "title": "Data Update hampering scores?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/415215",
  "author_name": "",
  "post_date": "2023-06-05T17:30:22.055860700Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I joined this competition again after a long hiatus so just wanted to check if anyone observed the same things I did.</p>\n<p>For reference, lets check the <a href=\"https://www.kaggle.com/code/nickcgray/gait-prediction?scriptVersionId=126602493\" target=\"_blank\">best public notebook</a>. It has a score of 0.310 but if you fork and run it right now, you'll get a ~0.28x score. This can also be seen for the <a href=\"https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets\" target=\"_blank\">second best notebook</a>.</p>\n<p>Now both of these were published before the data update. The same notebook can't perform the same with new data. That is fine but I find quite a few entries in the LB where it's clear they ran the notebook before the update (or saved those models) and their 0.310 scores are intact. So I am not sure how the rescoring impacted such cases.</p>\n<p>In my opinion, I think it will be great if we get access to the previous data too since it heavily favors the LB (at least the public one)</p>",
  "messages": [
    {
      "id": "2288808",
      "postDate": "06/05/2023 17:30:22",
      "content": "<p>I joined this competition again after a long hiatus so just wanted to check if anyone observed the same things I did.</p>\n<p>For reference, lets check the <a href=\"https://www.kaggle.com/code/nickcgray/gait-prediction?scriptVersionId=126602493\" target=\"_blank\">best public notebook</a>. It has a score of 0.310 but if you fork and run it right now, you'll get a ~0.28x score. This can also be seen for the <a href=\"https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets\" target=\"_blank\">second best notebook</a>.</p>\n<p>Now both of these were published before the data update. The same notebook can't perform the same with new data. That is fine but I find quite a few entries in the LB where it's clear they ran the notebook before the update (or saved those models) and their 0.310 scores are intact. So I am not sure how the rescoring impacted such cases.</p>\n<p>In my opinion, I think it will be great if we get access to the previous data too since it heavily favors the LB (at least the public one)</p>",
      "rawMarkdown": "I joined this competition again after a long hiatus so just wanted to check if anyone observed the same things I did.\n\nFor reference, lets check the [best public notebook](https://www.kaggle.com/code/nickcgray/gait-prediction?scriptVersionId=126602493). It has a score of 0.310 but if you fork and run it right now, you'll get a ~0.28x score. This can also be seen for the [second best notebook](https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets).\n\nNow both of these were published before the data update. The same notebook can't perform the same with new data. That is fine but I find quite a few entries in the LB where it's clear they ran the notebook before the update (or saved those models) and their 0.310 scores are intact. So I am not sure how the rescoring impacted such cases.\n\nIn my opinion, I think it will be great if we get access to the previous data too since it heavily favors the LB (at least the public one)",
      "votes": null
    },
    {
      "id": "2288953",
      "postDate": "06/05/2023 19:41:19",
      "content": "<p>A little late for that?</p>",
      "rawMarkdown": "A little late for that?",
      "votes": null
    },
    {
      "id": "2289101",
      "postDate": "06/05/2023 23:18:20",
      "content": "<p>Yes perhaps. Not sure why this was not discussed earlier.</p>",
      "rawMarkdown": "Yes perhaps. Not sure why this was not discussed earlier.",
      "votes": null
    },
    {
      "id": "2289286",
      "postDate": "06/06/2023 04:13:58",
      "content": "<p>Hey! I observed the same thing. Submitted a few notebooks before the data update, the score was 0.302 on public LB, started working on it again just a couple of days back again, and saw the same notebook gave 0.286. Even though my public LB showed 0.302, I had to work my way up again to 0.308 on updated data from 0.286 (because that's the true score) I am pretty sure that LB scores are quite mixed up and don't truly indicate the present notebooks standings. I've observed a couple of submissions with a score of 0.310 where last submission was months ago. It's odd how the LB scores were recalculated after the data update. </p>",
      "rawMarkdown": "Hey! I observed the same thing. Submitted a few notebooks before the data update, the score was 0.302 on public LB, started working on it again just a couple of days back again, and saw the same notebook gave 0.286. Even though my public LB showed 0.302, I had to work my way up again to 0.308 on updated data from 0.286 (because that's the true score) I am pretty sure that LB scores are quite mixed up and don't truly indicate the present notebooks standings. I've observed a couple of submissions with a score of 0.310 where last submission was months ago. It's odd how the LB scores were recalculated after the data update.",
      "votes": null
    },
    {
      "id": "2290547",
      "postDate": "06/06/2023 23:49:05",
      "content": "<p>Yes that's exactly what I observed. My guess is that they had saved the models after training and ran submission is a separate notebook in which case it will remain unaffected by the data update.</p>",
      "rawMarkdown": "Yes that's exactly what I observed. My guess is that they had saved the models after training and ran submission is a separate notebook in which case it will remain unaffected by the data update.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2288953,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "06/05/2023 19:41:19",
      "content": "<p>A little late for that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2289101,
          "author_name": "mayukh18",
          "author_url": "",
          "post_date": "06/05/2023 23:18:20",
          "content": "<p>Yes perhaps. Not sure why this was not discussed earlier.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2289286,
      "author_name": "ahsuna123",
      "author_url": "",
      "post_date": "06/06/2023 04:13:58",
      "content": "<p>Hey! I observed the same thing. Submitted a few notebooks before the data update, the score was 0.302 on public LB, started working on it again just a couple of days back again, and saw the same notebook gave 0.286. Even though my public LB showed 0.302, I had to work my way up again to 0.308 on updated data from 0.286 (because that's the true score) I am pretty sure that LB scores are quite mixed up and don't truly indicate the present notebooks standings. I've observed a couple of submissions with a score of 0.310 where last submission was months ago. It's odd how the LB scores were recalculated after the data update. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2290547,
          "author_name": "mayukh18",
          "author_url": "",
          "post_date": "06/06/2023 23:49:05",
          "content": "<p>Yes that's exactly what I observed. My guess is that they had saved the models after training and ran submission is a separate notebook in which case it will remain unaffected by the data update.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2288808": "I joined this competition again after a long hiatus so just wanted to check if anyone observed the same things I did.\n\nFor reference, lets check the [best public notebook](https://www.kaggle.com/code/nickcgray/gait-prediction?scriptVersionId=126602493). It has a score of 0.310 but if you fork and run it right now, you'll get a ~0.28x score. This can also be seen for the [second best notebook](https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets).\n\nNow both of these were published before the data update. The same notebook can't perform the same with new data. That is fine but I find quite a few entries in the LB where it's clear they ran the notebook before the update (or saved those models) and their 0.310 scores are intact. So I am not sure how the rescoring impacted such cases.\n\nIn my opinion, I think it will be great if we get access to the previous data too since it heavily favors the LB (at least the public one)",
    "2288953": "A little late for that?",
    "2289101": "Yes perhaps. Not sure why this was not discussed earlier.",
    "2289286": "Hey! I observed the same thing. Submitted a few notebooks before the data update, the score was 0.302 on public LB, started working on it again just a couple of days back again, and saw the same notebook gave 0.286. Even though my public LB showed 0.302, I had to work my way up again to 0.308 on updated data from 0.286 (because that's the true score) I am pretty sure that LB scores are quite mixed up and don't truly indicate the present notebooks standings. I've observed a couple of submissions with a score of 0.310 where last submission was months ago. It's odd how the LB scores were recalculated after the data update.",
    "2290547": "Yes that's exactly what I observed. My guess is that they had saved the models after training and ran submission is a separate notebook in which case it will remain unaffected by the data update."
  },
  "source": "meta"
}