{
  "id": 389848,
  "title": "How to make submission",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/389848",
  "author_name": "Julius Bohnen",
  "post_date": "2023-02-23T07:24:38.413000",
  "votes": 0,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm stuck creating a submission file. <br>\nI've trained 18 xgboost models, one for each question, and stored them in a list. Now i want to make a submission but i have no idea how (am new, sorry). To be precise, how do i select the right data to make a prediction with and how do i write the prediction in the submission file? Maybe someone can help? If you need more info feel free to ask me</p>",
  "messages": [
    {
      "id": 2156946,
      "postDate": "2023-02-23T16:21:20.190Z",
      "content": "<p>Hi, I'm also new here and trained in ways similar to you. Hope my submission code could help. I mainly follows Chris <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\" target=\"_blank\">notebook</a>. I 'm not so sure how to add codes by markdown :) lol</p>\n<pre><code>models = {}\nfor t in range(1,19):\n    if t&lt;=3: grp = '0-4'\n    elif t&lt;=13: grp = '5-12'\n    elif t&lt;=18: grp = '13-22'\n    clf = XGBClassifier()\n    clf.load_model(work_path + f'XGB_question_{t}.xgb')\n    models[f'{grp}_{t}'] = clf\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nbest_threshold = {0: 0.631, 1: 0.613, 2: 0.607, 3: 0.653, 4: 0.625, 5: 0.637, 6: 0.637, 7: 0.605, 8: 0.631, 9: 0.602, 10: 0.636, 11: 0.66, 12: 0.612, 13: 0.647, 14: 0.626, 15: 0.63, 16: 0.61, 17: 0.602}\ncount = 0\nFEATURES_QUESTION = np.loadtxt(work_path + 'FEATURES_QUESTION.txt',dtype=str, delimiter = ' ').tolist()\n\n\nfor (sample_submission, test) in iter_test:\n    df = feature_engineer(test)\n    #if count == 0:\n        #FEATURES = [c for c in df.columns if c != 'level_group']\n        #count += 1\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int(p.item()&gt;best_threshold[t-1])\n    env.predict(sample_submission)\n</code></pre>",
      "rawMarkdown": "Hi, I'm also new here and trained in ways similar to you. Hope my submission code could help. I mainly follows Chris [notebook](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676). I 'm not so sure how to add codes by markdown :) lol\n\n\n```\nmodels = {}\nfor t in range(1,19):\n    if t<=3: grp = '0-4'\n    elif t<=13: grp = '5-12'\n    elif t<=18: grp = '13-22'\n    clf = XGBClassifier()\n    clf.load_model(work_path + f'XGB_question_{t}.xgb')\n    models[f'{grp}_{t}'] = clf\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nbest_threshold = {0: 0.631, 1: 0.613, 2: 0.607, 3: 0.653, 4: 0.625, 5: 0.637, 6: 0.637, 7: 0.605, 8: 0.631, 9: 0.602, 10: 0.636, 11: 0.66, 12: 0.612, 13: 0.647, 14: 0.626, 15: 0.63, 16: 0.61, 17: 0.602}\ncount = 0\nFEATURES_QUESTION = np.loadtxt(work_path + 'FEATURES_QUESTION.txt',dtype=str, delimiter = ' ').tolist()\n\n\nfor (sample_submission, test) in iter_test:\n    df = feature_engineer(test)\n    #if count == 0:\n        #FEATURES = [c for c in df.columns if c != 'level_group']\n        #count += 1\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int(p.item()>best_threshold[t-1])\n    env.predict(sample_submission)\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2157075,
          "postDate": "2023-02-23T17:51:40.043Z",
          "content": "<p>Thanks! It does indeed look similar. I still get the error, though.. So i must do it wrong. Basically, i copied everything starting from \"import jo_wilder\" to the end. My models are already stored the same way you did it. Then I changed FEATURE_QUESTION to a list of the features my models are using. And I changed clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1] into clf.predict_proba(df[FEATURES_QUESTION].astype('float32'))[:,1], so that every  model uses the same features. Still doesnt work :/</p>",
          "rawMarkdown": "Thanks! It does indeed look similar. I still get the error, though.. So i must do it wrong. Basically, i copied everything starting from \"import jo_wilder\" to the end. My models are already stored the same way you did it. Then I changed FEATURE_QUESTION to a list of the features my models are using. And I changed clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1] into clf.predict_proba(df[FEATURES_QUESTION].astype('float32'))[:,1], so that every  model uses the same features. Still doesnt work :/"
        }
      ]
    },
    {
      "id": 2156338,
      "postDate": "2023-02-23T09:25:55.650Z",
      "content": "<p>Have a look at this great notebook by <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> - it takes some pretrained models saved in a kaggle dataset, defines a feature engineering function (done with polars instead of pandas!), and then cells 4-6 set up the API needed for the submission and then loops through the data for submission to perform the relevant feature engineering and predict using the pretrained models <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference</a></p>\n<p>There's also this notebook for a sample submission that shows how that API works! <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference</a></p>",
      "rawMarkdown": "Have a look at this great notebook by @carnozhao - it takes some pretrained models saved in a kaggle dataset, defines a feature engineering function (done with polars instead of pandas!), and then cells 4-6 set up the API needed for the submission and then loops through the data for submission to perform the relevant feature engineering and predict using the pretrained models https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\n\nThere's also this notebook for a sample submission that shows how that API works! https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference",
      "replies": [
        {
          "id": 2156605,
          "postDate": "2023-02-23T12:37:51.757Z",
          "content": "<p>Thanks for your reply! I tried doing it accordingly to the first notebook you provided but get an error. </p>\n<h2>\"You must call <code>predict()</code> successfully before you can continue with <code>iter_test()</code>:</h2>\n<p>TypeError                                 Traceback (most recent call last)<br>\n/tmp/ipykernel_25/3104309027.py in <br>\n      3 levels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}<br>\n      4 questions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}<br>\n----&gt; 5 for (sample_submission, test) in iter_test:<br>\n      6     target_level_group = level_groups_reverse[test.level_group.iloc[0]]<br>\n      7     df = feature_engineer(test)</p>\n<p>TypeError: cannot unpack non-iterable NoneType object </p>\n<p>I did it like this. <br>\nlevel_groups = [\"0-4\", \"5-12\", \"13-22\"]<br>\nlevel_groups_reverse = {'0-4': 0, '5-12': 1, '13-22': 2}<br>\nlevels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}<br>\nquestions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}<br>\nfor (sample_submission, test) in iter_test:<br>\n    target_level_group = level_groups_reverse[test.level_group.iloc[0]]<br>\n    df = add_features(data_engineer(test))</p>\n<pre><code>fold = 0\npreds = []\nfor q in range(*questions[level_groups[target_level_group]]):\n    model = models[q - 1]\n    feature_cols = model.feature_names_\n    pred = model.predict_proba(df[FEATURES].astype(np.float32))[0,1]\n    preds.append(int(pred &gt; 0.63))\n\nsample_submission[\"correct\"] = preds\n\nenv.predict(sample_submission)\"\n</code></pre>\n<p>where models is the list of 18 models, and FEATURES is a list of column names used for predictions. df is created in the same way i created the df for training </p>",
          "rawMarkdown": "Thanks for your reply! I tried doing it accordingly to the first notebook you provided but get an error. \n\n\"You must call `predict()` successfully before you can continue with `iter_test()`:\n---------------------------------------------------------------------------\nTypeError                                 Traceback (most recent call last)\n/tmp/ipykernel_25/3104309027.py in <module>\n      3 levels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}\n      4 questions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}\n----> 5 for (sample_submission, test) in iter_test:\n      6     target_level_group = level_groups_reverse[test.level_group.iloc[0]]\n      7     df = feature_engineer(test)\n\nTypeError: cannot unpack non-iterable NoneType object \n\n\nI did it like this. \nlevel_groups = [\"0-4\", \"5-12\", \"13-22\"]\nlevel_groups_reverse = {'0-4': 0, '5-12': 1, '13-22': 2}\nlevels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}\nquestions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}\nfor (sample_submission, test) in iter_test:\n    target_level_group = level_groups_reverse[test.level_group.iloc[0]]\n    df = add_features(data_engineer(test))\n    \n    fold = 0\n    preds = []\n    for q in range(*questions[level_groups[target_level_group]]):\n        model = models[q - 1]\n        feature_cols = model.feature_names_\n        pred = model.predict_proba(df[FEATURES].astype(np.float32))[0,1]\n        preds.append(int(pred > 0.63))\n\n    sample_submission[\"correct\"] = preds\n\n    env.predict(sample_submission)\"\n\nwhere models is the list of 18 models, and FEATURES is a list of column names used for predictions. df is created in the same way i created the df for training ",
          "votes": 1,
          "replies": [
            {
              "id": 2157044,
              "postDate": "2023-02-23T17:37:40.007Z",
              "content": "<p>It looks like you might have already started iterating over iter_test. Restarting your notebook kernel and rerunning it should work with the code that you've got there - essentially that env.predict line has to run before the API can move on to the next piece of data</p>",
              "rawMarkdown": "It looks like you might have already started iterating over iter_test. Restarting your notebook kernel and rerunning it should work with the code that you've got there - essentially that env.predict line has to run before the API can move on to the next piece of data"
            },
            {
              "id": 2158729,
              "postDate": "2023-02-25T05:28:16.910Z",
              "content": "<p>Ah, that was indeed the problem. I thought I could just restart the run of the cell.. Thanks a lot!</p>",
              "rawMarkdown": "Ah, that was indeed the problem. I thought I could just restart the run of the cell.. Thanks a lot!"
            }
          ]
        }
      ]
    },
    {
      "id": 2156236,
      "postDate": "2023-02-23T07:24:38.413Z",
      "content": "<p>I'm stuck creating a submission file. <br>\nI've trained 18 xgboost models, one for each question, and stored them in a list. Now i want to make a submission but i have no idea how (am new, sorry). To be precise, how do i select the right data to make a prediction with and how do i write the prediction in the submission file? Maybe someone can help? If you need more info feel free to ask me</p>",
      "rawMarkdown": "I'm stuck creating a submission file. \nI've trained 18 xgboost models, one for each question, and stored them in a list. Now i want to make a submission but i have no idea how (am new, sorry). To be precise, how do i select the right data to make a prediction with and how do i write the prediction in the submission file? Maybe someone can help? If you need more info feel free to ask me\n"
    },
    {
      "id": 2156922,
      "postDate": "2023-02-23T16:06:29.427Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2156946,
      "author_name": "Moonlit",
      "author_url": "",
      "post_date": "2023-02-23T16:21:20.190000",
      "content": "<p>Hi, I'm also new here and trained in ways similar to you. Hope my submission code could help. I mainly follows Chris <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676\" target=\"_blank\">notebook</a>. I 'm not so sure how to add codes by markdown :) lol</p>\n<pre><code>models = {}\nfor t in range(1,19):\n    if t&lt;=3: grp = '0-4'\n    elif t&lt;=13: grp = '5-12'\n    elif t&lt;=18: grp = '13-22'\n    clf = XGBClassifier()\n    clf.load_model(work_path + f'XGB_question_{t}.xgb')\n    models[f'{grp}_{t}'] = clf\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nbest_threshold = {0: 0.631, 1: 0.613, 2: 0.607, 3: 0.653, 4: 0.625, 5: 0.637, 6: 0.637, 7: 0.605, 8: 0.631, 9: 0.602, 10: 0.636, 11: 0.66, 12: 0.612, 13: 0.647, 14: 0.626, 15: 0.63, 16: 0.61, 17: 0.602}\ncount = 0\nFEATURES_QUESTION = np.loadtxt(work_path + 'FEATURES_QUESTION.txt',dtype=str, delimiter = ' ').tolist()\n\n\nfor (sample_submission, test) in iter_test:\n    df = feature_engineer(test)\n    #if count == 0:\n        #FEATURES = [c for c in df.columns if c != 'level_group']\n        #count += 1\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int(p.item()&gt;best_threshold[t-1])\n    env.predict(sample_submission)\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2157075,
          "author_name": "Julius Bohnen",
          "author_url": "",
          "post_date": "2023-02-23T17:51:40.043000",
          "content": "<p>Thanks! It does indeed look similar. I still get the error, though.. So i must do it wrong. Basically, i copied everything starting from \"import jo_wilder\" to the end. My models are already stored the same way you did it. Then I changed FEATURE_QUESTION to a list of the features my models are using. And I changed clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1] into clf.predict_proba(df[FEATURES_QUESTION].astype('float32'))[:,1], so that every  model uses the same features. Still doesnt work :/</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2156338,
      "author_name": "Jude Hunt",
      "author_url": "",
      "post_date": "2023-02-23T09:25:55.650000",
      "content": "<p>Have a look at this great notebook by <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> - it takes some pretrained models saved in a kaggle dataset, defines a feature engineering function (done with polars instead of pandas!), and then cells 4-6 set up the API needed for the submission and then loops through the data for submission to perform the relevant feature engineering and predict using the pretrained models <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference</a></p>\n<p>There's also this notebook for a sample submission that shows how that API works! <a href=\"https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\" target=\"_blank\">https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2156605,
          "author_name": "Julius Bohnen",
          "author_url": "",
          "post_date": "2023-02-23T12:37:51.757000",
          "content": "<p>Thanks for your reply! I tried doing it accordingly to the first notebook you provided but get an error. </p>\n<h2>\"You must call <code>predict()</code> successfully before you can continue with <code>iter_test()</code>:</h2>\n<p>TypeError                                 Traceback (most recent call last)<br>\n/tmp/ipykernel_25/3104309027.py in <br>\n      3 levels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}<br>\n      4 questions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}<br>\n----&gt; 5 for (sample_submission, test) in iter_test:<br>\n      6     target_level_group = level_groups_reverse[test.level_group.iloc[0]]<br>\n      7     df = feature_engineer(test)</p>\n<p>TypeError: cannot unpack non-iterable NoneType object </p>\n<p>I did it like this. <br>\nlevel_groups = [\"0-4\", \"5-12\", \"13-22\"]<br>\nlevel_groups_reverse = {'0-4': 0, '5-12': 1, '13-22': 2}<br>\nlevels = {'0-4': (0, 5), '5-12': (5, 13), '13-22': (13, 23)}<br>\nquestions = {'0-4': (1, 4), '5-12': (4, 14), '13-22': (14, 19)}<br>\nfor (sample_submission, test) in iter_test:<br>\n    target_level_group = level_groups_reverse[test.level_group.iloc[0]]<br>\n    df = add_features(data_engineer(test))</p>\n<pre><code>fold = 0\npreds = []\nfor q in range(*questions[level_groups[target_level_group]]):\n    model = models[q - 1]\n    feature_cols = model.feature_names_\n    pred = model.predict_proba(df[FEATURES].astype(np.float32))[0,1]\n    preds.append(int(pred &gt; 0.63))\n\nsample_submission[\"correct\"] = preds\n\nenv.predict(sample_submission)\"\n</code></pre>\n<p>where models is the list of 18 models, and FEATURES is a list of column names used for predictions. df is created in the same way i created the df for training </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2157044,
              "author_name": "Jude Hunt",
              "author_url": "",
              "post_date": "2023-02-23T17:37:40.007000",
              "content": "<p>It looks like you might have already started iterating over iter_test. Restarting your notebook kernel and rerunning it should work with the code that you've got there - essentially that env.predict line has to run before the API can move on to the next piece of data</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2158729,
              "author_name": "Julius Bohnen",
              "author_url": "",
              "post_date": "2023-02-25T05:28:16.910000",
              "content": "<p>Ah, that was indeed the problem. I thought I could just restart the run of the cell.. Thanks a lot!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2156922,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-23T16:06:29.427000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2156946": "Hi, I'm also new here and trained in ways similar to you. Hope my submission code could help. I mainly follows Chris [notebook](https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-676). I 'm not so sure how to add codes by markdown :) lol\n\n\n```\nmodels = {}\nfor t in range(1,19):\n    if t<=3: grp = '0-4'\n    elif t<=13: grp = '5-12'\n    elif t<=18: grp = '13-22'\n    clf = XGBClassifier()\n    clf.load_model(work_path + f'XGB_question_{t}.xgb')\n    models[f'{grp}_{t}'] = clf\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nbest_threshold = {0: 0.631, 1: 0.613, 2: 0.607, 3: 0.653, 4: 0.625, 5: 0.637, 6: 0.637, 7: 0.605, 8: 0.631, 9: 0.602, 10: 0.636, 11: 0.66, 12: 0.612, 13: 0.647, 14: 0.626, 15: 0.63, 16: 0.61, 17: 0.602}\ncount = 0\nFEATURES_QUESTION = np.loadtxt(work_path + 'FEATURES_QUESTION.txt',dtype=str, delimiter = ' ').tolist()\n\n\nfor (sample_submission, test) in iter_test:\n    df = feature_engineer(test)\n    #if count == 0:\n        #FEATURES = [c for c in df.columns if c != 'level_group']\n        #count += 1\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES_QUESTION[t-1]].astype('float32'))[:,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int(p.item()>best_threshold[t-1])\n    env.predict(sample_submission)\n```",
    "2156338": "Have a look at this great notebook by @carnozhao - it takes some pretrained models saved in a kaggle dataset, defines a feature engineering function (done with polars instead of pandas!), and then cells 4-6 set up the API needed for the submission and then loops through the data for submission to perform the relevant feature engineering and predict using the pretrained models https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference\n\nThere's also this notebook for a sample submission that shows how that API works! https://www.kaggle.com/code/carnozhao/cpu-catboost-baseline-using-polars-inference",
    "2156236": "I'm stuck creating a submission file. \nI've trained 18 xgboost models, one for each question, and stored them in a list. Now i want to make a submission but i have no idea how (am new, sorry). To be precise, how do i select the right data to make a prediction with and how do i write the prediction in the submission file? Maybe someone can help? If you need more info feel free to ask me\n",
    "2156922": ""
  }
}