{
  "id": 415988,
  "title": "What is the difference between public and private Dataset?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/415988",
  "author_name": "",
  "post_date": "2023-06-09T01:45:23.154277Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulation to all of the winners😄</p>\n<p>Our team had a shakedown(11→69th), just as we had feared.🤣<br>\nHowever, I learned many things, such as feature design techniques and data visualization techniques.<br>\nAnd most importantly, I enjoyed Kaggling. 👍<br>\nI would like to thank the competition organizers.</p>\n<p>By the Way,<br>\nFrom the middle of the competition, it was assumed that the Public Dataset could be biased data.<br>\nCould you tell us what data was actually present in Public and Private?<br>\nI would like to learn more about this competition, if you don't mind, so that I can reflect on this competition and make the most of the next time✌️</p>\n<p>Also, if you were changing the bias of the data between Public and Private, could you tell us what your intention was in designing the distribution of the data as the competition organizer?</p>",
  "messages": [
    {
      "id": "2293172",
      "postDate": "06/09/2023 01:45:23",
      "content": "<p>Congratulation to all of the winners😄</p>\n<p>Our team had a shakedown(11→69th), just as we had feared.🤣<br>\nHowever, I learned many things, such as feature design techniques and data visualization techniques.<br>\nAnd most importantly, I enjoyed Kaggling. 👍<br>\nI would like to thank the competition organizers.</p>\n<p>By the Way,<br>\nFrom the middle of the competition, it was assumed that the Public Dataset could be biased data.<br>\nCould you tell us what data was actually present in Public and Private?<br>\nI would like to learn more about this competition, if you don't mind, so that I can reflect on this competition and make the most of the next time✌️</p>\n<p>Also, if you were changing the bias of the data between Public and Private, could you tell us what your intention was in designing the distribution of the data as the competition organizer?</p>",
      "rawMarkdown": "Congratulation to all of the winners😄\n\nOur team had a shakedown(11→69th), just as we had feared.🤣\nHowever, I learned many things, such as feature design techniques and data visualization techniques.\nAnd most importantly, I enjoyed Kaggling. 👍\nI would like to thank the competition organizers.\n\nBy the Way,\nFrom the middle of the competition, it was assumed that the Public Dataset could be biased data.\nCould you tell us what data was actually present in Public and Private?\nI would like to learn more about this competition, if you don't mind, so that I can reflect on this competition and make the most of the next time✌️\n\nAlso, if you were changing the bias of the data between Public and Private, could you tell us what your intention was in designing the distribution of the data as the competition organizer?",
      "votes": null
    },
    {
      "id": "2293173",
      "postDate": "06/09/2023 01:58:19",
      "content": "<p>I have the same curiosity. My local cv(around 0.32-0.34) matchs my private score, but public score gets higher than 0.4.</p>",
      "rawMarkdown": "I have the same curiosity. My local cv(around 0.32-0.34) matchs my private score, but public score gets higher than 0.4.",
      "votes": null
    },
    {
      "id": "2293984",
      "postDate": "06/09/2023 16:20:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hidebu\" target=\"_blank\">@hidebu</a>, As mentioned in the dataset description <code>the train / public test / private test splits were formed by grouping on subjects and stratifying on datasets</code>. There was no other intentional bias in the split design. Whatever differences in distribution were present were the result of the random sampling.</p>",
      "rawMarkdown": "Hi @hidebu, As mentioned in the dataset description `the train / public test / private test splits were formed by grouping on subjects and stratifying on datasets`. There was no other intentional bias in the split design. Whatever differences in distribution were present were the result of the random sampling.",
      "votes": null
    },
    {
      "id": "2294316",
      "postDate": "06/10/2023 00:57:56",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nThank you reply kindly. <br>\nI understood it to be the result of differences in the probability of developing FOG in each patient.</p>",
      "rawMarkdown": "ryanholbrook \nThank you reply kindly. \nI understood it to be the result of differences in the probability of developing FOG in each patient.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2293173,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "06/09/2023 01:58:19",
      "content": "<p>I have the same curiosity. My local cv(around 0.32-0.34) matchs my private score, but public score gets higher than 0.4.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2293984,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "06/09/2023 16:20:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hidebu\" target=\"_blank\">@hidebu</a>, As mentioned in the dataset description <code>the train / public test / private test splits were formed by grouping on subjects and stratifying on datasets</code>. There was no other intentional bias in the split design. Whatever differences in distribution were present were the result of the random sampling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2294316,
          "author_name": "hidebu",
          "author_url": "",
          "post_date": "06/10/2023 00:57:56",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nThank you reply kindly. <br>\nI understood it to be the result of differences in the probability of developing FOG in each patient.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2293172": "Congratulation to all of the winners😄\n\nOur team had a shakedown(11→69th), just as we had feared.🤣\nHowever, I learned many things, such as feature design techniques and data visualization techniques.\nAnd most importantly, I enjoyed Kaggling. 👍\nI would like to thank the competition organizers.\n\nBy the Way,\nFrom the middle of the competition, it was assumed that the Public Dataset could be biased data.\nCould you tell us what data was actually present in Public and Private?\nI would like to learn more about this competition, if you don't mind, so that I can reflect on this competition and make the most of the next time✌️\n\nAlso, if you were changing the bias of the data between Public and Private, could you tell us what your intention was in designing the distribution of the data as the competition organizer?",
    "2293173": "I have the same curiosity. My local cv(around 0.32-0.34) matchs my private score, but public score gets higher than 0.4.",
    "2293984": "Hi @hidebu, As mentioned in the dataset description `the train / public test / private test splits were formed by grouping on subjects and stratifying on datasets`. There was no other intentional bias in the split design. Whatever differences in distribution were present were the result of the random sampling.",
    "2294316": "ryanholbrook \nThank you reply kindly. \nI understood it to be the result of differences in the probability of developing FOG in each patient."
  },
  "source": "meta"
}