{
  "id": 393228,
  "title": "NFolds Inference runtime limit",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/393228",
  "author_name": "",
  "post_date": "2023-03-08T14:21:16.510420Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I use the Nfold model to make predictions that exceed the time limit. I wonder if everyone has used the Nfold model to make predictions?</p>\n<p>🤕</p>",
  "messages": [
    {
      "id": "2173631",
      "postDate": "03/08/2023 14:21:16",
      "content": "<p>I use the Nfold model to make predictions that exceed the time limit. I wonder if everyone has used the Nfold model to make predictions?</p>\n<p>🤕</p>",
      "rawMarkdown": "I use the Nfold model to make predictions that exceed the time limit. I wonder if everyone has used the Nfold model to make predictions?\n\n🤕",
      "votes": null
    },
    {
      "id": "2173766",
      "postDate": "03/08/2023 16:12:33",
      "content": "<p>I used 5fold models. For me, the bottleneck of time cost is feature engineering instead of the Nfold prediction.</p>",
      "rawMarkdown": "I used 5fold models. For me, the bottleneck of time cost is feature engineering instead of the Nfold prediction.",
      "votes": null
    },
    {
      "id": "2174061",
      "postDate": "03/08/2023 20:12:19",
      "content": "<p>I used 5 folds as well, but I had the same issue. I would recommend retraining a single model on all of the data using the parameters from your Nfold models. For example, if you're training an XGB with early stopping, you could take the average of the <code>best_ntree_limit</code> across the folds for each question and retrain one model (per question) on all the data, using that average <code>best_ntree_limit</code>. </p>\n<p>It could also be worth making a dummy prediction where you run your feature engineering but make a prediction of all zeros. That should tell you how long your feature engineering takes compared to your model predictions.</p>",
      "rawMarkdown": "I used 5 folds as well, but I had the same issue. I would recommend retraining a single model on all of the data using the parameters from your Nfold models. For example, if you're training an XGB with early stopping, you could take the average of the `best_ntree_limit` across the folds for each question and retrain one model (per question) on all the data, using that average `best_ntree_limit`. \n\nIt could also be worth making a dummy prediction where you run your feature engineering but make a prediction of all zeros. That should tell you how long your feature engineering takes compared to your model predictions.",
      "votes": null
    },
    {
      "id": "2174096",
      "postDate": "03/08/2023 20:38:30",
      "content": "<p>if you want to use n-fold prediction, make sure you load n models into mem before first iteration: reading model from file on every iteration is a bottleneck</p>",
      "rawMarkdown": "if you want to use n-fold prediction, make sure you load n models into mem before first iteration: reading model from file on every iteration is a bottleneck",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2173766,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "03/08/2023 16:12:33",
      "content": "<p>I used 5fold models. For me, the bottleneck of time cost is feature engineering instead of the Nfold prediction.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2174061,
      "author_name": "judehunt23",
      "author_url": "",
      "post_date": "03/08/2023 20:12:19",
      "content": "<p>I used 5 folds as well, but I had the same issue. I would recommend retraining a single model on all of the data using the parameters from your Nfold models. For example, if you're training an XGB with early stopping, you could take the average of the <code>best_ntree_limit</code> across the folds for each question and retrain one model (per question) on all the data, using that average <code>best_ntree_limit</code>. </p>\n<p>It could also be worth making a dummy prediction where you run your feature engineering but make a prediction of all zeros. That should tell you how long your feature engineering takes compared to your model predictions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2174096,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "03/08/2023 20:38:30",
      "content": "<p>if you want to use n-fold prediction, make sure you load n models into mem before first iteration: reading model from file on every iteration is a bottleneck</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2173631": "I use the Nfold model to make predictions that exceed the time limit. I wonder if everyone has used the Nfold model to make predictions?\n\n🤕",
    "2173766": "I used 5fold models. For me, the bottleneck of time cost is feature engineering instead of the Nfold prediction.",
    "2174061": "I used 5 folds as well, but I had the same issue. I would recommend retraining a single model on all of the data using the parameters from your Nfold models. For example, if you're training an XGB with early stopping, you could take the average of the `best_ntree_limit` across the folds for each question and retrain one model (per question) on all the data, using that average `best_ntree_limit`. \n\nIt could also be worth making a dummy prediction where you run your feature engineering but make a prediction of all zeros. That should tell you how long your feature engineering takes compared to your model predictions.",
    "2174096": "if you want to use n-fold prediction, make sure you load n models into mem before first iteration: reading model from file on every iteration is a bottleneck"
  },
  "source": "meta"
}