{
  "id": 496416,
  "title": "Testing takes an extremely long time. Ways to speed it up?",
  "url": "/competitions/leash-BELKA/discussion/496416",
  "author_name": "",
  "post_date": "2024-04-21T04:27:34.297091Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm noticing that the testing process on Kaggle cloud takes an insanely long time (1hr + for a simple random forest method). I'm assuming it is purely due to the size of the test set, as training (30,000 samples) was pretty quick.</p>\n<p>Does anyone have any suggestions for speeding this up?</p>\n<p>Am I better to do this locally where I have access to more CPU cores?</p>",
  "messages": [
    {
      "id": "2765264",
      "postDate": "04/21/2024 04:27:34",
      "content": "<p>I'm noticing that the testing process on Kaggle cloud takes an insanely long time (1hr + for a simple random forest method). I'm assuming it is purely due to the size of the test set, as training (30,000 samples) was pretty quick.</p>\n<p>Does anyone have any suggestions for speeding this up?</p>\n<p>Am I better to do this locally where I have access to more CPU cores?</p>",
      "rawMarkdown": "I'm noticing that the testing process on Kaggle cloud takes an insanely long time (1hr + for a simple random forest method). I'm assuming it is purely due to the size of the test set, as training (30,000 samples) was pretty quick.\n\nDoes anyone have any suggestions for speeding this up?\n\nAm I better to do this locally where I have access to more CPU cores?",
      "votes": null
    },
    {
      "id": "2765340",
      "postDate": "04/21/2024 05:27:13",
      "content": "<p>Of course, this is not a coding competition, so you are free to do your process locally and then submit the parquet file to Kaggle separately without using a Kaggle kernel at all <a href=\"https://www.kaggle.com/benfield\" target=\"_blank\">@benfield</a> </p>",
      "rawMarkdown": "Of course, this is not a coding competition, so you are free to do your process locally and then submit the parquet file to Kaggle separately without using a Kaggle kernel at all @benfield",
      "votes": null
    },
    {
      "id": "2765380",
      "postDate": "04/21/2024 05:50:50",
      "content": "<p>check my other post:<br>\nfor submission of test (1674896 SMILES)<br>\nxgboost 1xgpu : less than one min <br>\ndeep learning transformer (molformer), 2xgpu: 5 min </p>\n<p>all time exclude feature extraction (e.g. rdkits ecfp, tokenization which are less than 30 min on multicore cpu) </p>",
      "rawMarkdown": "check my other post:\nfor submission of test (1674896 SMILES)\nxgboost 1xgpu : less than one min \ndeep learning transformer (molformer), 2xgpu: 5 min \n\nall time exclude feature extraction (e.g. rdkits ecfp, tokenization which are less than 30 min on multicore cpu)",
      "votes": null
    },
    {
      "id": "2766410",
      "postDate": "04/21/2024 17:27:02",
      "content": "<p>1 hour seems long. Unless you're using a for loop?</p>\n<p>Oh, if you're using this notebook or a variation:<br>\n<a href=\"https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\" target=\"_blank\">https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest</a></p>\n<p>The issue isn't making the predictions, the issue is getting the fingerprint feature. That taking close to an hour for 1.6 million rows seems reasonable.</p>\n<p>You can speed it up by:</p>\n<ul>\n<li>Running once per molecule instead of once per row. (2x speed up)</li>\n<li>Running once in notebook A and saving the result. Then using the result in notebook B. Now you only pay the 30 minute cost each time you ever change your fingerprint features algorithm.</li>\n</ul>",
      "rawMarkdown": "1 hour seems long. Unless you're using a for loop?\n\nOh, if you're using this notebook or a variation:\nhttps://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\n\nThe issue isn't making the predictions, the issue is getting the fingerprint feature. That taking close to an hour for 1.6 million rows seems reasonable.\n\nYou can speed it up by:\n* Running once per molecule instead of once per row. (2x speed up)\n* Running once in notebook A and saving the result. Then using the result in notebook B. Now you only pay the 30 minute cost each time you ever change your fingerprint features algorithm.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2765340,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/21/2024 05:27:13",
      "content": "<p>Of course, this is not a coding competition, so you are free to do your process locally and then submit the parquet file to Kaggle separately without using a Kaggle kernel at all <a href=\"https://www.kaggle.com/benfield\" target=\"_blank\">@benfield</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2765380,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/21/2024 05:50:50",
      "content": "<p>check my other post:<br>\nfor submission of test (1674896 SMILES)<br>\nxgboost 1xgpu : less than one min <br>\ndeep learning transformer (molformer), 2xgpu: 5 min </p>\n<p>all time exclude feature extraction (e.g. rdkits ecfp, tokenization which are less than 30 min on multicore cpu) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2766410,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/21/2024 17:27:02",
      "content": "<p>1 hour seems long. Unless you're using a for loop?</p>\n<p>Oh, if you're using this notebook or a variation:<br>\n<a href=\"https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\" target=\"_blank\">https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest</a></p>\n<p>The issue isn't making the predictions, the issue is getting the fingerprint feature. That taking close to an hour for 1.6 million rows seems reasonable.</p>\n<p>You can speed it up by:</p>\n<ul>\n<li>Running once per molecule instead of once per row. (2x speed up)</li>\n<li>Running once in notebook A and saving the result. Then using the result in notebook B. Now you only pay the 30 minute cost each time you ever change your fingerprint features algorithm.</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2765264": "I'm noticing that the testing process on Kaggle cloud takes an insanely long time (1hr + for a simple random forest method). I'm assuming it is purely due to the size of the test set, as training (30,000 samples) was pretty quick.\n\nDoes anyone have any suggestions for speeding this up?\n\nAm I better to do this locally where I have access to more CPU cores?",
    "2765340": "Of course, this is not a coding competition, so you are free to do your process locally and then submit the parquet file to Kaggle separately without using a Kaggle kernel at all @benfield",
    "2765380": "check my other post:\nfor submission of test (1674896 SMILES)\nxgboost 1xgpu : less than one min \ndeep learning transformer (molformer), 2xgpu: 5 min \n\nall time exclude feature extraction (e.g. rdkits ecfp, tokenization which are less than 30 min on multicore cpu)",
    "2766410": "1 hour seems long. Unless you're using a for loop?\n\nOh, if you're using this notebook or a variation:\nhttps://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\n\nThe issue isn't making the predictions, the issue is getting the fingerprint feature. That taking close to an hour for 1.6 million rows seems reasonable.\n\nYou can speed it up by:\n* Running once per molecule instead of once per row. (2x speed up)\n* Running once in notebook A and saving the result. Then using the result in notebook B. Now you only pay the 30 minute cost each time you ever change your fingerprint features algorithm."
  },
  "source": "meta"
}