{
  "id": 52178,
  "title": "Best sample size and how much we can push the kaggle servers?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52178",
  "author_name": "",
  "post_date": "2018-03-17T01:02:40.534382300Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My train sample has around 20 million records. The kernel has been running for more than 3 hours. What is the running time limit set by kaggle before the kernel gets killed?\nCan the fellow Kagglers in this competition post their sample size and the kernel running durations?\n This will help us in deciding the best sample size and how much we can push the kaggle servers?</p>",
  "messages": [
    {
      "id": "297389",
      "postDate": "03/17/2018 01:02:40",
      "content": "<p>My train sample has around 20 million records. The kernel has been running for more than 3 hours. What is the running time limit set by kaggle before the kernel gets killed?\nCan the fellow Kagglers in this competition post their sample size and the kernel running durations?\n This will help us in deciding the best sample size and how much we can push the kaggle servers?</p>",
      "rawMarkdown": "My train sample has around 20 million records. The kernel has been running for more than 3 hours. What is the running time limit set by kaggle before the kernel gets killed?\nCan the fellow Kagglers in this competition post their sample size and the kernel running durations?\n This will help us in deciding the best sample size and how much we can push the kaggle servers?",
      "votes": null
    },
    {
      "id": "297401",
      "postDate": "03/17/2018 01:48:09",
      "content": "<p>Which model are you using?</p>\n\n<p>I have tested <strong><a href=\"https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-auc-0-9787/code\">Single lightGBM</a></strong>  just now with following configs: </p>\n\n<ul>\n<li>sample: first 40 million rows </li>\n<li>train: first 37 million rows from sample</li>\n<li>validation : last 3 million rows from sample</li>\n<li>train's auc: 0.977824    valid's auc: 0.978751</li>\n<li>Public Lb: 0.9671</li>\n<li>run time: ~51 mins on Kaggle (with 4 threads)</li>\n</ul>\n\n<p>Maximum threads are 32 but speed is best with 4 threads (I have played around with it a lot)</p>",
      "rawMarkdown": "Which model are you using?\n\nI have tested **[Single lightGBM][1]**  just now with following configs: \n\n - sample: first 40 million rows \n - train: first 37 million rows from sample\n - validation : last 3 million rows from sample\n - train's auc: 0.977824\tvalid's auc: 0.978751\n - Public Lb: 0.9671\n - run time: ~51 mins on Kaggle (with 4 threads)\n\nMaximum threads are 32 but speed is best with 4 threads (I have played around with it a lot)\n\n  [1]: https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-auc-0-9787/code",
      "votes": null
    },
    {
      "id": "297407",
      "postDate": "03/17/2018 02:15:41",
      "content": "<p>Thanks for the response. I am using XGBoost . Playing  around with feature engineering . 4 hours and it’s still running . </p>",
      "rawMarkdown": "Thanks for the response. I am using XGBoost . Playing  around with feature engineering . 4 hours and it’s still running .",
      "votes": null
    },
    {
      "id": "297413",
      "postDate": "03/17/2018 02:43:27",
      "content": "<p>My pleasure! Try XGBoost on hist mode. It's super fast and gives very good accuracy. You may want to check out this <a href=\"https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-9639?scriptVersionId=2815141/code\">excellent XGBoost</a>  kernel which uses 55 million rows and has run time just ~14 mins</p>",
      "rawMarkdown": "My pleasure! Try XGBoost on hist mode. It's super fast and gives very good accuracy. You may want to check out this [excellent XGBoost][1]  kernel which uses 55 million rows and has run time just ~14 mins\n\n\n  [1]: https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-9639?scriptVersionId=2815141/code",
      "votes": null
    },
    {
      "id": "297417",
      "postDate": "03/17/2018 03:05:05",
      "content": "<p>Wow .. let me check that out ..thanks!!!</p>",
      "rawMarkdown": "Wow .. let me check that out ..thanks!!!",
      "votes": null
    },
    {
      "id": "299351",
      "postDate": "03/20/2018 20:10:50",
      "content": "<p>Created a new train sample and used your code. Jumped to 43rd places. Thanks a lot, man!!</p>",
      "rawMarkdown": "Created a new train sample and used your code. Jumped to 43rd places. Thanks a lot, man!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 297401,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "03/17/2018 01:48:09",
      "content": "<p>Which model are you using?</p>\n\n<p>I have tested <strong><a href=\"https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-auc-0-9787/code\">Single lightGBM</a></strong>  just now with following configs: </p>\n\n<ul>\n<li>sample: first 40 million rows </li>\n<li>train: first 37 million rows from sample</li>\n<li>validation : last 3 million rows from sample</li>\n<li>train's auc: 0.977824    valid's auc: 0.978751</li>\n<li>Public Lb: 0.9671</li>\n<li>run time: ~51 mins on Kaggle (with 4 threads)</li>\n</ul>\n\n<p>Maximum threads are 32 but speed is best with 4 threads (I have played around with it a lot)</p>",
      "votes": null,
      "replies": [
        {
          "id": 297407,
          "author_name": "gopisaran",
          "author_url": "",
          "post_date": "03/17/2018 02:15:41",
          "content": "<p>Thanks for the response. I am using XGBoost . Playing  around with feature engineering . 4 hours and it’s still running . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297413,
          "author_name": "pranav84",
          "author_url": "",
          "post_date": "03/17/2018 02:43:27",
          "content": "<p>My pleasure! Try XGBoost on hist mode. It's super fast and gives very good accuracy. You may want to check out this <a href=\"https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-9639?scriptVersionId=2815141/code\">excellent XGBoost</a>  kernel which uses 55 million rows and has run time just ~14 mins</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297417,
          "author_name": "gopisaran",
          "author_url": "",
          "post_date": "03/17/2018 03:05:05",
          "content": "<p>Wow .. let me check that out ..thanks!!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299351,
          "author_name": "gopisaran",
          "author_url": "",
          "post_date": "03/20/2018 20:10:50",
          "content": "<p>Created a new train sample and used your code. Jumped to 43rd places. Thanks a lot, man!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "297389": "My train sample has around 20 million records. The kernel has been running for more than 3 hours. What is the running time limit set by kaggle before the kernel gets killed?\nCan the fellow Kagglers in this competition post their sample size and the kernel running durations?\n This will help us in deciding the best sample size and how much we can push the kaggle servers?",
    "297401": "Which model are you using?\n\nI have tested **[Single lightGBM][1]**  just now with following configs: \n\n - sample: first 40 million rows \n - train: first 37 million rows from sample\n - validation : last 3 million rows from sample\n - train's auc: 0.977824\tvalid's auc: 0.978751\n - Public Lb: 0.9671\n - run time: ~51 mins on Kaggle (with 4 threads)\n\nMaximum threads are 32 but speed is best with 4 threads (I have played around with it a lot)\n\n  [1]: https://www.kaggle.com/pranav84/lightgbm-fixing-unbalanced-data-auc-0-9787/code",
    "297407": "Thanks for the response. I am using XGBoost . Playing  around with feature engineering . 4 hours and it’s still running .",
    "297413": "My pleasure! Try XGBoost on hist mode. It's super fast and gives very good accuracy. You may want to check out this [excellent XGBoost][1]  kernel which uses 55 million rows and has run time just ~14 mins\n\n\n  [1]: https://www.kaggle.com/joaopmpeinado/single-xgboost-lb-0-9639?scriptVersionId=2815141/code",
    "297417": "Wow .. let me check that out ..thanks!!!",
    "299351": "Created a new train sample and used your code. Jumped to 43rd places. Thanks a lot, man!!"
  },
  "source": "meta"
}