{
  "id": 198069,
  "title": "Curious the performance about the more data",
  "url": "/competitions/riiid-test-answer-prediction/discussion/198069",
  "author_name": "",
  "post_date": "2020-11-19T15:58:23.696631100Z",
  "votes": 19,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Because my computer RAM limited（only 16G）,Now my use 5M training data(60 features),it seems that my features didn't import my model perference,just curious if use more data can import the performance so that  i need to rent a server.</p>",
  "messages": [
    {
      "id": "1083996",
      "postDate": "11/19/2020 15:58:23",
      "content": "<p>Because my computer RAM limited（only 16G）,Now my use 5M training data(60 features),it seems that my features didn't import my model perference,just curious if use more data can import the performance so that  i need to rent a server.</p>",
      "rawMarkdown": "Because my computer RAM limited（only 16G）,Now my use 5M training data(60 features),it seems that my features didn't import my model perference,just curious if use more data can import the performance so that  i need to rent a server.",
      "votes": null
    },
    {
      "id": "1084024",
      "postDate": "11/19/2020 16:21:29",
      "content": "<p>In my case it helps ~0.01 moving from 10M to 80M rows.</p>",
      "rawMarkdown": "In my case it helps ~0.01 moving from 10M to 80M rows.",
      "votes": null
    },
    {
      "id": "1084029",
      "postDate": "11/19/2020 16:28:49",
      "content": "<p>Amazing,thanks for your information</p>",
      "rawMarkdown": "Amazing,thanks for your information",
      "votes": null
    },
    {
      "id": "1084030",
      "postDate": "11/19/2020 16:29:09",
      "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> you mean you add more training data and your score boost from 0.787 to 0.797?👀</p>",
      "rawMarkdown": "bacicnikola you mean you add more training data and your score boost from 0.787 to 0.797?👀",
      "votes": null
    },
    {
      "id": "1084033",
      "postDate": "11/19/2020 16:34:33",
      "content": "<p>5M rows and #12 is amazing! You can use free 300$ of GCP.</p>",
      "rawMarkdown": "5M rows and #12 is amazing! You can use free 300$ of GCP.",
      "votes": null
    },
    {
      "id": "1084053",
      "postDate": "11/19/2020 17:10:20",
      "content": "<p>Well actually it boost my validation score. I'm always making a submission on full dataset, so I don't actually know the LB difference…</p>",
      "rawMarkdown": "Well actually it boost my validation score. I'm always making a submission on full dataset, so I don't actually know the LB difference...",
      "votes": null
    },
    {
      "id": "1084190",
      "postDate": "11/19/2020 20:25:23",
      "content": "<p>My CV is correlated with LB. <br>\n20M rows to 80M rows also boost by +0.01 like <a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a></p>",
      "rawMarkdown": "My CV is correlated with LB. \n20M rows to 80M rows also boost by +0.01 like @bacicnikola",
      "votes": null
    },
    {
      "id": "1084193",
      "postDate": "11/19/2020 20:28:41",
      "content": "<p>by \"making a submission on full dataset\" do you mean you use the most recent historical records from full dataset in test time (e.g. to do all the merge etc.), or do you mean retraining a classifier on the full dataset (number if boosting rounds selected from CV)?</p>",
      "rawMarkdown": "by \"making a submission on full dataset\" do you mean you use the most recent historical records from full dataset in test time (e.g. to do all the merge etc.), or do you mean retraining a classifier on the full dataset (number if boosting rounds selected from CV)?",
      "votes": null
    },
    {
      "id": "1084447",
      "postDate": "11/20/2020 03:32:38",
      "content": "<p>Do you submit 1 model to predict?<br>\ni submit 10 models, which training 10m row. data is spilted by mod(user_id, 10)</p>",
      "rawMarkdown": "Do you submit 1 model to predict?\ni submit 10 models, which training 10m row. data is spilted by mod(user_id, 10)",
      "votes": null
    },
    {
      "id": "1084449",
      "postDate": "11/20/2020 03:36:20",
      "content": "<p>yes,only one lgb</p>",
      "rawMarkdown": "yes,only one lgb",
      "votes": null
    },
    {
      "id": "1084453",
      "postDate": "11/20/2020 03:43:22",
      "content": "<p>Great job,thanks</p>",
      "rawMarkdown": "Great job,thanks",
      "votes": null
    },
    {
      "id": "1084537",
      "postDate": "11/20/2020 06:04:39",
      "content": "<p>Amazing, 791 with only 5M data!!!</p>",
      "rawMarkdown": "Amazing, 791 with only 5M data!!!",
      "votes": null
    },
    {
      "id": "1084667",
      "postDate": "11/20/2020 09:07:56",
      "content": "<p>I mean retraining a classifier on the full dataset.</p>",
      "rawMarkdown": "I mean retraining a classifier on the full dataset.",
      "votes": null
    },
    {
      "id": "1084755",
      "postDate": "11/20/2020 11:02:58",
      "content": "<p>Good to hear that! I'm currently using even less than 10M rows for training, so hopefully that big improvement is awaiting me, too :)</p>",
      "rawMarkdown": "Good to hear that! I'm currently using even less than 10M rows for training, so hopefully that big improvement is awaiting me, too :)",
      "votes": null
    },
    {
      "id": "1085560",
      "postDate": "11/21/2020 01:20:33",
      "content": "<p>As the data grows, the size of the state variable increases rapidly. As a result, when using the full amount of data, there is an out-of-memory situation for online infering. This is what happened to me. I don't know anyone else.</p>",
      "rawMarkdown": "As the data grows, the size of the state variable increases rapidly. As a result, when using the full amount of data, there is an out-of-memory situation for online infering. This is what happened to me. I don't know anyone else.",
      "votes": null
    },
    {
      "id": "1085857",
      "postDate": "11/21/2020 08:43:38",
      "content": "<p>Yes, me too, I'm fighting to avoid OOM. The more features the more memory used.</p>",
      "rawMarkdown": "Yes, me too, I'm fighting to avoid OOM. The more features the more memory used.",
      "votes": null
    },
    {
      "id": "1086924",
      "postDate": "11/22/2020 07:19:04",
      "content": "<p>When you say \"my computer RAM\" - if you are talking about your local machine than it's not a real hard chore to increase the effective RAM if you have some space on an SSD in the system.</p>\n<p>On Ubuntu Linux you can use SSD memory as swapfile - I add an SSD to my system and use 200GB of that as swap.  <br>\nOn that same machine when I boot up in Windows 10 I use 200GB for virtual memory.  (if you guess the SSD I add is a half GB than you win the prize):</p>\n<p>The Linux is a bit tricky - email me and I can send you the code snippets to get the swapfile raised from the default (think it's 2GB) to a larger value.</p>\n<p>This could also work if you have standard hard drive rather than SSD's but the speed performance on a hard drive really sucks.  With SSD's on Ubuntu I can feel the bump in performance when the swap is needed but it's something you can live with.</p>\n<p>I have this setup running on 4 different local machines - you can still have kernels die with what are likely memory issues but for the most part it works.</p>",
      "rawMarkdown": "When you say \"my computer RAM\" - if you are talking about your local machine than it's not a real hard chore to increase the effective RAM if you have some space on an SSD in the system.\n\nOn Ubuntu Linux you can use SSD memory as swapfile - I add an SSD to my system and use 200GB of that as swap.  \nOn that same machine when I boot up in Windows 10 I use 200GB for virtual memory.  (if you guess the SSD I add is a half GB than you win the prize):\n\nThe Linux is a bit tricky - email me and I can send you the code snippets to get the swapfile raised from the default (think it's 2GB) to a larger value.\n\nThis could also work if you have standard hard drive rather than SSD's but the speed performance on a hard drive really sucks.  With SSD's on Ubuntu I can feel the bump in performance when the swap is needed but it's something you can live with.\n\nI have this setup running on 4 different local machines - you can still have kernels die with what are likely memory issues but for the most part it works.",
      "votes": null
    },
    {
      "id": "1091433",
      "postDate": "11/26/2020 02:46:34",
      "content": "<p>thanks for you help</p>",
      "rawMarkdown": "thanks for you help",
      "votes": null
    },
    {
      "id": "1100722",
      "postDate": "12/03/2020 09:21:08",
      "content": "<p>Amazing! I can use my 32G RAM joining!</p>",
      "rawMarkdown": "Amazing! I can use my 32G RAM joining!",
      "votes": null
    },
    {
      "id": "1101545",
      "postDate": "12/04/2020 02:18:28",
      "content": "<p>You mean use 80M to training or just generating features？</p>",
      "rawMarkdown": "You mean use 80M to training or just generating features？",
      "votes": null
    },
    {
      "id": "1108680",
      "postDate": "12/10/2020 22:33:53",
      "content": "<p>My CV used to be correlated with LB really well, but not in this case of training more data…Training on the full dataset improved my CV by ~0.01 (like others have mentioned) but didn't improve my LB as much.  <br>\nNot sure if anyone else is facing the same situation.  </p>",
      "rawMarkdown": "My CV used to be correlated with LB really well, but not in this case of training more data...Training on the full dataset improved my CV by ~0.01 (like others have mentioned) but didn't improve my LB as much.  \nNot sure if anyone else is facing the same situation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1084024,
      "author_name": "bacicnikola",
      "author_url": "",
      "post_date": "11/19/2020 16:21:29",
      "content": "<p>In my case it helps ~0.01 moving from 10M to 80M rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1084029,
          "author_name": "mahluo",
          "author_url": "",
          "post_date": "11/19/2020 16:28:49",
          "content": "<p>Amazing,thanks for your information</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084030,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "11/19/2020 16:29:09",
          "content": "<p><a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a> you mean you add more training data and your score boost from 0.787 to 0.797?👀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084053,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "11/19/2020 17:10:20",
          "content": "<p>Well actually it boost my validation score. I'm always making a submission on full dataset, so I don't actually know the LB difference…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084193,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "11/19/2020 20:28:41",
          "content": "<p>by \"making a submission on full dataset\" do you mean you use the most recent historical records from full dataset in test time (e.g. to do all the merge etc.), or do you mean retraining a classifier on the full dataset (number if boosting rounds selected from CV)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084667,
          "author_name": "bacicnikola",
          "author_url": "",
          "post_date": "11/20/2020 09:07:56",
          "content": "<p>I mean retraining a classifier on the full dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084755,
          "author_name": "alijs1",
          "author_url": "",
          "post_date": "11/20/2020 11:02:58",
          "content": "<p>Good to hear that! I'm currently using even less than 10M rows for training, so hopefully that big improvement is awaiting me, too :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1084033,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "11/19/2020 16:34:33",
      "content": "<p>5M rows and #12 is amazing! You can use free 300$ of GCP.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1084190,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "11/19/2020 20:25:23",
      "content": "<p>My CV is correlated with LB. <br>\n20M rows to 80M rows also boost by +0.01 like <a href=\"https://www.kaggle.com/bacicnikola\" target=\"_blank\">@bacicnikola</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1084453,
          "author_name": "mahluo",
          "author_url": "",
          "post_date": "11/20/2020 03:43:22",
          "content": "<p>Great job,thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1084447,
      "author_name": "kurupical",
      "author_url": "",
      "post_date": "11/20/2020 03:32:38",
      "content": "<p>Do you submit 1 model to predict?<br>\ni submit 10 models, which training 10m row. data is spilted by mod(user_id, 10)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1084449,
          "author_name": "mahluo",
          "author_url": "",
          "post_date": "11/20/2020 03:36:20",
          "content": "<p>yes,only one lgb</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1084537,
      "author_name": "shichenyang",
      "author_url": "",
      "post_date": "11/20/2020 06:04:39",
      "content": "<p>Amazing, 791 with only 5M data!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1085560,
      "author_name": "shichenyang",
      "author_url": "",
      "post_date": "11/21/2020 01:20:33",
      "content": "<p>As the data grows, the size of the state variable increases rapidly. As a result, when using the full amount of data, there is an out-of-memory situation for online infering. This is what happened to me. I don't know anyone else.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1085857,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "11/21/2020 08:43:38",
          "content": "<p>Yes, me too, I'm fighting to avoid OOM. The more features the more memory used.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1086924,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/22/2020 07:19:04",
      "content": "<p>When you say \"my computer RAM\" - if you are talking about your local machine than it's not a real hard chore to increase the effective RAM if you have some space on an SSD in the system.</p>\n<p>On Ubuntu Linux you can use SSD memory as swapfile - I add an SSD to my system and use 200GB of that as swap.  <br>\nOn that same machine when I boot up in Windows 10 I use 200GB for virtual memory.  (if you guess the SSD I add is a half GB than you win the prize):</p>\n<p>The Linux is a bit tricky - email me and I can send you the code snippets to get the swapfile raised from the default (think it's 2GB) to a larger value.</p>\n<p>This could also work if you have standard hard drive rather than SSD's but the speed performance on a hard drive really sucks.  With SSD's on Ubuntu I can feel the bump in performance when the swap is needed but it's something you can live with.</p>\n<p>I have this setup running on 4 different local machines - you can still have kernels die with what are likely memory issues but for the most part it works.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091433,
          "author_name": "mahluo",
          "author_url": "",
          "post_date": "11/26/2020 02:46:34",
          "content": "<p>thanks for you help</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100722,
      "author_name": "daishu",
      "author_url": "",
      "post_date": "12/03/2020 09:21:08",
      "content": "<p>Amazing! I can use my 32G RAM joining!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1101545,
      "author_name": "juzqyxs",
      "author_url": "",
      "post_date": "12/04/2020 02:18:28",
      "content": "<p>You mean use 80M to training or just generating features？</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1108680,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "12/10/2020 22:33:53",
      "content": "<p>My CV used to be correlated with LB really well, but not in this case of training more data…Training on the full dataset improved my CV by ~0.01 (like others have mentioned) but didn't improve my LB as much.  <br>\nNot sure if anyone else is facing the same situation.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1083996": "Because my computer RAM limited（only 16G）,Now my use 5M training data(60 features),it seems that my features didn't import my model perference,just curious if use more data can import the performance so that  i need to rent a server.",
    "1084024": "In my case it helps ~0.01 moving from 10M to 80M rows.",
    "1084029": "Amazing,thanks for your information",
    "1084030": "bacicnikola you mean you add more training data and your score boost from 0.787 to 0.797?👀",
    "1084033": "5M rows and #12 is amazing! You can use free 300$ of GCP.",
    "1084053": "Well actually it boost my validation score. I'm always making a submission on full dataset, so I don't actually know the LB difference...",
    "1084190": "My CV is correlated with LB. \n20M rows to 80M rows also boost by +0.01 like @bacicnikola",
    "1084193": "by \"making a submission on full dataset\" do you mean you use the most recent historical records from full dataset in test time (e.g. to do all the merge etc.), or do you mean retraining a classifier on the full dataset (number if boosting rounds selected from CV)?",
    "1084447": "Do you submit 1 model to predict?\ni submit 10 models, which training 10m row. data is spilted by mod(user_id, 10)",
    "1084449": "yes,only one lgb",
    "1084453": "Great job,thanks",
    "1084537": "Amazing, 791 with only 5M data!!!",
    "1084667": "I mean retraining a classifier on the full dataset.",
    "1084755": "Good to hear that! I'm currently using even less than 10M rows for training, so hopefully that big improvement is awaiting me, too :)",
    "1085560": "As the data grows, the size of the state variable increases rapidly. As a result, when using the full amount of data, there is an out-of-memory situation for online infering. This is what happened to me. I don't know anyone else.",
    "1085857": "Yes, me too, I'm fighting to avoid OOM. The more features the more memory used.",
    "1086924": "When you say \"my computer RAM\" - if you are talking about your local machine than it's not a real hard chore to increase the effective RAM if you have some space on an SSD in the system.\n\nOn Ubuntu Linux you can use SSD memory as swapfile - I add an SSD to my system and use 200GB of that as swap.  \nOn that same machine when I boot up in Windows 10 I use 200GB for virtual memory.  (if you guess the SSD I add is a half GB than you win the prize):\n\nThe Linux is a bit tricky - email me and I can send you the code snippets to get the swapfile raised from the default (think it's 2GB) to a larger value.\n\nThis could also work if you have standard hard drive rather than SSD's but the speed performance on a hard drive really sucks.  With SSD's on Ubuntu I can feel the bump in performance when the swap is needed but it's something you can live with.\n\nI have this setup running on 4 different local machines - you can still have kernels die with what are likely memory issues but for the most part it works.",
    "1091433": "thanks for you help",
    "1100722": "Amazing! I can use my 32G RAM joining!",
    "1101545": "You mean use 80M to training or just generating features？",
    "1108680": "My CV used to be correlated with LB really well, but not in this case of training more data...Training on the full dataset improved my CV by ~0.01 (like others have mentioned) but didn't improve my LB as much.  \nNot sure if anyone else is facing the same situation."
  },
  "source": "meta"
}