{
  "id": 415027,
  "title": "Seeking a help/advise for the difference during Scoring",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/415027",
  "author_name": "",
  "post_date": "2023-06-04T18:53:09.859103700Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all , sorry that i new to kaggle and the question may looks stupied . <br>\ni am refering Chris ' code (<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680)\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680)</a>, and my code is with 2 parts (1 trainning, 1 inference )<br>\nPrevious , the score is very close to my F1 score in training part , <br>\nBut recently i am trying to predict the result by accumulating the grp 0-4 's data into grp 5-12,  and also accumulating grp 0-4 , 5-12 for grp 13-22,  the scoring is quiet different between the trainning and infer part . (score is 0.698 in training , but the LB score is only 0.655) <br>\nCould anyone help to give me some help or suggestion ?  Thanks a lot <br>\nMy submission code is below :<br>\n`</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n\nbest_threshold = \n\ntestAccumulated = pd.DataFrame()\n\n\n\n (test,sample_submission)  iter_test:\n\n    grp = test.level_group.values[]\n    session_id = test.session_id.values[]\n    testAccumulated = pd.concat([testAccumulated, test], axis=)\n\n    df = (pl.from_pandas(testAccumulated.loc[testAccumulated[] == session_id])\n          .drop([, , ])\n          .with_columns(columns))\n\n    \n\n    test_predict = feature_engineer_pl(df, grp, use_extra=, feature_suffix=)\n    (, grp)\n    (, session_id)\n\n\n\n\n\n\n\n    \n    a,b = limits[grp]\n    (, a,b)\n     t  (a, b):\n        FEATURES = importance_dict[(t)]\n\n        model = XGBClassifier()\n        model.load_model()\n\n        p = model.predict_proba(test_predict[FEATURES].astype())[:,]\n        ((p))\n        mask = sample_submission.session_id..contains()\n        sample_submission.loc[mask,] = (p &gt; best_threshold).astype()\n\n    env.predict(sample_submission[[, ]])\n</code></pre>\n<p>`</p>",
  "messages": [
    {
      "id": "2287667",
      "postDate": "06/04/2023 18:53:09",
      "content": "<p>Hi all , sorry that i new to kaggle and the question may looks stupied . <br>\ni am refering Chris ' code (<a href=\"https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680)\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680)</a>, and my code is with 2 parts (1 trainning, 1 inference )<br>\nPrevious , the score is very close to my F1 score in training part , <br>\nBut recently i am trying to predict the result by accumulating the grp 0-4 's data into grp 5-12,  and also accumulating grp 0-4 , 5-12 for grp 13-22,  the scoring is quiet different between the trainning and infer part . (score is 0.698 in training , but the LB score is only 0.655) <br>\nCould anyone help to give me some help or suggestion ?  Thanks a lot <br>\nMy submission code is below :<br>\n`</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n\nbest_threshold = \n\ntestAccumulated = pd.DataFrame()\n\n\n\n (test,sample_submission)  iter_test:\n\n    grp = test.level_group.values[]\n    session_id = test.session_id.values[]\n    testAccumulated = pd.concat([testAccumulated, test], axis=)\n\n    df = (pl.from_pandas(testAccumulated.loc[testAccumulated[] == session_id])\n          .drop([, , ])\n          .with_columns(columns))\n\n    \n\n    test_predict = feature_engineer_pl(df, grp, use_extra=, feature_suffix=)\n    (, grp)\n    (, session_id)\n\n\n\n\n\n\n\n    \n    a,b = limits[grp]\n    (, a,b)\n     t  (a, b):\n        FEATURES = importance_dict[(t)]\n\n        model = XGBClassifier()\n        model.load_model()\n\n        p = model.predict_proba(test_predict[FEATURES].astype())[:,]\n        ((p))\n        mask = sample_submission.session_id..contains()\n        sample_submission.loc[mask,] = (p &gt; best_threshold).astype()\n\n    env.predict(sample_submission[[, ]])\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "Hi all , sorry that i new to kaggle and the question may looks stupied . \ni am refering Chris ' code (https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680), and my code is with 2 parts (1 trainning, 1 inference )\nPrevious , the score is very close to my F1 score in training part , \nBut recently i am trying to predict the result by accumulating the grp 0-4 's data into grp 5-12,  and also accumulating grp 0-4 , 5-12 for grp 13-22,  the scoring is quiet different between the trainning and infer part . (score is 0.698 in training , but the LB score is only 0.655) \nCould anyone help to give me some help or suggestion ?  Thanks a lot \nMy submission code is below :\n`\n\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nbest_threshold = 0.625\n\ntestAccumulated = pd.DataFrame()\n\n# historical_meta = defaultdict(list)\n\nfor (test,sample_submission) in iter_test:\n    \n    grp = test.level_group.values[0]\n    session_id = test.session_id.values[0]\n    testAccumulated = pd.concat([testAccumulated, test], axis=0)\n\n    df = (pl.from_pandas(testAccumulated.loc[testAccumulated[\"session_id\"] == session_id])\n          .drop([\"fullscreen\", \"hq\", \"music\"])\n          .with_columns(columns))\n    \n    # FEATURE ENGINEER TEST DATA\n\n    test_predict = feature_engineer_pl(df, grp, use_extra=True, feature_suffix='')\n    print('grp:', grp)\n    print('session_id:', session_id)\n#     if grp == '0-4':\n#         df = df1\n#     elif grp == '5-12': \n#         df = df2\n#     elif grp == '13-22': \n#         df = df3\n    \n    # INFER TEST DATA\n    a,b = limits[grp]\n    print('a,b:', a,b)\n    for t in range(a, b):\n        FEATURES = importance_dict[str(t)]\n\n        model = XGBClassifier()\n        model.load_model(f'/kaggle/input/modelsaved/XGB_question{t}.xgb')\n\n        p = model.predict_proba(test_predict[FEATURES].astype('float32'))[:,1]\n        print((p))\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = (p > best_threshold).astype(int)\n            \n    env.predict(sample_submission[['session_id', 'correct']])\n```\n`",
      "votes": null
    },
    {
      "id": "2297156",
      "postDate": "06/12/2023 11:47:45",
      "content": "<p>Following <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">Jack's Thread</a>, I found that for these models adding:</p>\n<p><code>test = test.sort_values('index').reset_index(drop=True)</code></p>\n<p>right after the loop helps.</p>",
      "rawMarkdown": "Following [Jack's Thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512), I found that for these models adding:\n\n`test = test.sort_values('index').reset_index(drop=True)`\n\nright after the loop helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2297156,
      "author_name": "danielphalen",
      "author_url": "",
      "post_date": "06/12/2023 11:47:45",
      "content": "<p>Following <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">Jack's Thread</a>, I found that for these models adding:</p>\n<p><code>test = test.sort_values('index').reset_index(drop=True)</code></p>\n<p>right after the loop helps.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2287667": "Hi all , sorry that i new to kaggle and the question may looks stupied . \ni am refering Chris ' code (https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680), and my code is with 2 parts (1 trainning, 1 inference )\nPrevious , the score is very close to my F1 score in training part , \nBut recently i am trying to predict the result by accumulating the grp 0-4 's data into grp 5-12,  and also accumulating grp 0-4 , 5-12 for grp 13-22,  the scoring is quiet different between the trainning and infer part . (score is 0.698 in training , but the LB score is only 0.655) \nCould anyone help to give me some help or suggestion ?  Thanks a lot \nMy submission code is below :\n`\n\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nbest_threshold = 0.625\n\ntestAccumulated = pd.DataFrame()\n\n# historical_meta = defaultdict(list)\n\nfor (test,sample_submission) in iter_test:\n    \n    grp = test.level_group.values[0]\n    session_id = test.session_id.values[0]\n    testAccumulated = pd.concat([testAccumulated, test], axis=0)\n\n    df = (pl.from_pandas(testAccumulated.loc[testAccumulated[\"session_id\"] == session_id])\n          .drop([\"fullscreen\", \"hq\", \"music\"])\n          .with_columns(columns))\n    \n    # FEATURE ENGINEER TEST DATA\n\n    test_predict = feature_engineer_pl(df, grp, use_extra=True, feature_suffix='')\n    print('grp:', grp)\n    print('session_id:', session_id)\n#     if grp == '0-4':\n#         df = df1\n#     elif grp == '5-12': \n#         df = df2\n#     elif grp == '13-22': \n#         df = df3\n    \n    # INFER TEST DATA\n    a,b = limits[grp]\n    print('a,b:', a,b)\n    for t in range(a, b):\n        FEATURES = importance_dict[str(t)]\n\n        model = XGBClassifier()\n        model.load_model(f'/kaggle/input/modelsaved/XGB_question{t}.xgb')\n\n        p = model.predict_proba(test_predict[FEATURES].astype('float32'))[:,1]\n        print((p))\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = (p > best_threshold).astype(int)\n            \n    env.predict(sample_submission[['session_id', 'correct']])\n```\n`",
    "2297156": "Following [Jack's Thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512), I found that for these models adding:\n\n`test = test.sort_values('index').reset_index(drop=True)`\n\nright after the loop helps."
  },
  "source": "meta"
}