{
  "id": 203450,
  "title": "Can Brute force win this competition?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/203450",
  "author_name": "",
  "post_date": "2020-12-15T10:40:30.726300100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I don't have much access to any resources, and I saw that people getting 0.78 scores are also using around 15 M rows. But are the top position people using the entire data and getting such high scores? Has anyone seen any significant difference in using the entire data rather than using a subset of it?</p>",
  "messages": [
    {
      "id": "1113268",
      "postDate": "12/15/2020 10:40:30",
      "content": "<p>I don't have much access to any resources, and I saw that people getting 0.78 scores are also using around 15 M rows. But are the top position people using the entire data and getting such high scores? Has anyone seen any significant difference in using the entire data rather than using a subset of it?</p>",
      "rawMarkdown": "I don't have much access to any resources, and I saw that people getting 0.78 scores are also using around 15 M rows. But are the top position people using the entire data and getting such high scores? Has anyone seen any significant difference in using the entire data rather than using a subset of it?",
      "votes": null
    },
    {
      "id": "1113298",
      "postDate": "12/15/2020 11:05:17",
      "content": "<p>You can achieve .~.79 with just 8-10M rows pretty much as what i have seen folks reporting here on forums, so i would say, jump in and don't worry about compute, you always have GCP's credits etc if you want to go that way!</p>",
      "rawMarkdown": "You can achieve .~.79 with just 8-10M rows pretty much as what i have seen folks reporting here on forums, so i would say, jump in and don't worry about compute, you always have GCP's credits etc if you want to go that way!",
      "votes": null
    },
    {
      "id": "1127601",
      "postDate": "12/26/2020 17:12:04",
      "content": "<p>Very often I get higher score with less train data… I am also wondering, is it required to use the entire dataset for a final submission? Or can we use less data if the score is higher…. </p>",
      "rawMarkdown": "Very often I get higher score with less train data... I am also wondering, is it required to use the entire dataset for a final submission? Or can we use less data if the score is higher....",
      "votes": null
    },
    {
      "id": "1127642",
      "postDate": "12/26/2020 17:52:57",
      "content": "<p>more data ,higher score, but the diff &lt;0.01,<br>\nI train 26M ,47features,just use kaggle kernel,</p>",
      "rawMarkdown": "more data ,higher score, but the diff <0.01,\nI train 26M ,47features,just use kaggle kernel,",
      "votes": null
    },
    {
      "id": "1137336",
      "postDate": "01/03/2021 21:06:35",
      "content": "<p>surely, if not overfitting, more training data leads to better LB</p>",
      "rawMarkdown": "surely, if not overfitting, more training data leads to better LB",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1113298,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "12/15/2020 11:05:17",
      "content": "<p>You can achieve .~.79 with just 8-10M rows pretty much as what i have seen folks reporting here on forums, so i would say, jump in and don't worry about compute, you always have GCP's credits etc if you want to go that way!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1127601,
      "author_name": "fjodorfomin",
      "author_url": "",
      "post_date": "12/26/2020 17:12:04",
      "content": "<p>Very often I get higher score with less train data… I am also wondering, is it required to use the entire dataset for a final submission? Or can we use less data if the score is higher…. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1127642,
      "author_name": "yangxiaoshuai",
      "author_url": "",
      "post_date": "12/26/2020 17:52:57",
      "content": "<p>more data ,higher score, but the diff &lt;0.01,<br>\nI train 26M ,47features,just use kaggle kernel,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1137336,
      "author_name": "elvinagammed",
      "author_url": "",
      "post_date": "01/03/2021 21:06:35",
      "content": "<p>surely, if not overfitting, more training data leads to better LB</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1113268": "I don't have much access to any resources, and I saw that people getting 0.78 scores are also using around 15 M rows. But are the top position people using the entire data and getting such high scores? Has anyone seen any significant difference in using the entire data rather than using a subset of it?",
    "1113298": "You can achieve .~.79 with just 8-10M rows pretty much as what i have seen folks reporting here on forums, so i would say, jump in and don't worry about compute, you always have GCP's credits etc if you want to go that way!",
    "1127601": "Very often I get higher score with less train data... I am also wondering, is it required to use the entire dataset for a final submission? Or can we use less data if the score is higher....",
    "1127642": "more data ,higher score, but the diff <0.01,\nI train 26M ,47features,just use kaggle kernel,",
    "1137336": "surely, if not overfitting, more training data leads to better LB"
  },
  "source": "meta"
}