{
  "id": 401294,
  "title": "How to use decision tree, random forest to generate prediction results fulfilling the submission requirements?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/401294",
  "author_name": "",
  "post_date": "2023-04-12T16:04:44.073906300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello bosses, We're a group of year 1 CS student. Can any bosses teach us to use random forest or decision tree over the test dataset to generate prediction results fulfilling submission requirements? We're a bit confused what to do next with all those models already generated some ACCURACY but we don't know how to apply them over test datasets.</p>\n<p>Your opinion and help will be greatly appreciated !</p>",
  "messages": [
    {
      "id": "2219500",
      "postDate": "04/12/2023 16:04:44",
      "content": "<p>Hello bosses, We're a group of year 1 CS student. Can any bosses teach us to use random forest or decision tree over the test dataset to generate prediction results fulfilling submission requirements? We're a bit confused what to do next with all those models already generated some ACCURACY but we don't know how to apply them over test datasets.</p>\n<p>Your opinion and help will be greatly appreciated !</p>",
      "rawMarkdown": "Hello bosses, We're a group of year 1 CS student. Can any bosses teach us to use random forest or decision tree over the test dataset to generate prediction results fulfilling submission requirements? We're a bit confused what to do next with all those models already generated some ACCURACY but we don't know how to apply them over test datasets.\n\nYour opinion and help will be greatly appreciated !",
      "votes": null
    },
    {
      "id": "2219508",
      "postDate": "04/12/2023 16:09:48",
      "content": "<p>Welcome! I published a random forest baseline <a href=\"https://www.kaggle.com/code/cdeotte/random-forest-baseline-0-664\" target=\"_blank\">here</a>. This notebook was written to train with the original data and now throws error on new larger train data, so you will need to use the template of my updated XGB notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\" target=\"_blank\">here</a> and then use the random forest model. For example, you will need to update the following (and can copy the code from my XGB notebook)</p>\n<ul>\n<li><p>load and feature engineer train data in 10 chunks to avoid memory error</p></li>\n<li><p>change prediction loop to use </p>\n<pre><code>p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\nmask = sample_submission.session_id.str.contains(f'q{t}')\nsample_submission.loc[mask,'correct'] = int( p &gt; best_threshold )\n</code></pre></li>\n</ul>\n<p>Instead of <code>[:,1]</code> and <code>p.item()</code> because the sample API has 3 users and submit API has 1 user per for-loop. Enjoy!</p>",
      "rawMarkdown": "Welcome! I published a random forest baseline [here][1]. This notebook was written to train with the original data and now throws error on new larger train data, so you will need to use the template of my updated XGB notebook [here][2] and then use the random forest model. For example, you will need to update the following (and can copy the code from my XGB notebook)\n* load and feature engineer train data in 10 chunks to avoid memory error\n* change prediction loop to use \n\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n\nInstead of `[:,1]` and `p.item()` because the sample API has 3 users and submit API has 1 user per for-loop. Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/random-forest-baseline-0-664\n[2]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680",
      "votes": null
    },
    {
      "id": "2219840",
      "postDate": "04/12/2023 23:46:07",
      "content": "<p>i have a notebook, its using random forest, you can have a look: <a href=\"https://www.kaggle.com/code/tyeestudio/feature-correlation-and-importance\" target=\"_blank\">https://www.kaggle.com/code/tyeestudio/feature-correlation-and-importance</a></p>\n<p>note: generating submission.csv works fine in this notebook, but submit to leaderboard will fail ( will fix it, but currently focusing other stuff )</p>",
      "rawMarkdown": "i have a notebook, its using random forest, you can have a look: https://www.kaggle.com/code/tyeestudio/feature-correlation-and-importance\n\nnote: generating submission.csv works fine in this notebook, but submit to leaderboard will fail ( will fix it, but currently focusing other stuff )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2219508,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/12/2023 16:09:48",
      "content": "<p>Welcome! I published a random forest baseline <a href=\"https://www.kaggle.com/code/cdeotte/random-forest-baseline-0-664\" target=\"_blank\">here</a>. This notebook was written to train with the original data and now throws error on new larger train data, so you will need to use the template of my updated XGB notebook <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680\" target=\"_blank\">here</a> and then use the random forest model. For example, you will need to update the following (and can copy the code from my XGB notebook)</p>\n<ul>\n<li><p>load and feature engineer train data in 10 chunks to avoid memory error</p></li>\n<li><p>change prediction loop to use </p>\n<pre><code>p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\nmask = sample_submission.session_id.str.contains(f'q{t}')\nsample_submission.loc[mask,'correct'] = int( p &gt; best_threshold )\n</code></pre></li>\n</ul>\n<p>Instead of <code>[:,1]</code> and <code>p.item()</code> because the sample API has 3 users and submit API has 1 user per for-loop. Enjoy!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2219840,
      "author_name": "tyeestudio",
      "author_url": "",
      "post_date": "04/12/2023 23:46:07",
      "content": "<p>i have a notebook, its using random forest, you can have a look: <a href=\"https://www.kaggle.com/code/tyeestudio/feature-correlation-and-importance\" target=\"_blank\">https://www.kaggle.com/code/tyeestudio/feature-correlation-and-importance</a></p>\n<p>note: generating submission.csv works fine in this notebook, but submit to leaderboard will fail ( will fix it, but currently focusing other stuff )</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2219500": "Hello bosses, We're a group of year 1 CS student. Can any bosses teach us to use random forest or decision tree over the test dataset to generate prediction results fulfilling submission requirements? We're a bit confused what to do next with all those models already generated some ACCURACY but we don't know how to apply them over test datasets.\n\nYour opinion and help will be greatly appreciated !",
    "2219508": "Welcome! I published a random forest baseline [here][1]. This notebook was written to train with the original data and now throws error on new larger train data, so you will need to use the template of my updated XGB notebook [here][2] and then use the random forest model. For example, you will need to update the following (and can copy the code from my XGB notebook)\n* load and feature engineer train data in 10 chunks to avoid memory error\n* change prediction loop to use \n\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n\nInstead of `[:,1]` and `p.item()` because the sample API has 3 users and submit API has 1 user per for-loop. Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/random-forest-baseline-0-664\n[2]: https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680",
    "2219840": "i have a notebook, its using random forest, you can have a look: https://www.kaggle.com/code/tyeestudio/feature-correlation-and-importance\n\nnote: generating submission.csv works fine in this notebook, but submit to leaderboard will fail ( will fix it, but currently focusing other stuff )"
  },
  "source": "meta"
}