{
  "id": 402640,
  "title": "My model always predict that all score of the test data is 1.",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/402640",
  "author_name": "",
  "post_date": "2023-04-19T06:34:31.091956600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>My model always predict that all score of the test data is 1. However my model seems to work on train data. Then I have no idea about how I can solve this problem.</p>\n<pre><code> jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {:(,), :(,), :(,)}\n\n (test, sample_submission)  iter_test:\n    df = pd.read_csv(, index_col=[])\n\n    \n    grp = test.level_group.values[]\n    a,b = limits[grp]\n     t  (a,b):\n        clf = models[]\n        p = clf.predict_proba(df[FEATURES].astype())[,]\n        mask = sample_submission.session_id..contains()\n        sample_submission.loc[mask,] = ( p &gt; best_threshold )\n\n    env.predict(sample_submission)\n</code></pre>\n<p>It might be a key to solve that this code raise an warning<br>\n<code>This version of the API is not optimized and should not be used to estimate the runtime of your code on the hidden test set.</code></p>\n<p>I use kaggle notebook.</p>\n<p>if you need more information, please feel free to let me know.</p>",
  "messages": [
    {
      "id": "2226704",
      "postDate": "04/19/2023 06:34:31",
      "content": "<p>My model always predict that all score of the test data is 1. However my model seems to work on train data. Then I have no idea about how I can solve this problem.</p>\n<pre><code> jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {:(,), :(,), :(,)}\n\n (test, sample_submission)  iter_test:\n    df = pd.read_csv(, index_col=[])\n\n    \n    grp = test.level_group.values[]\n    a,b = limits[grp]\n     t  (a,b):\n        clf = models[]\n        p = clf.predict_proba(df[FEATURES].astype())[,]\n        mask = sample_submission.session_id..contains()\n        sample_submission.loc[mask,] = ( p &gt; best_threshold )\n\n    env.predict(sample_submission)\n</code></pre>\n<p>It might be a key to solve that this code raise an warning<br>\n<code>This version of the API is not optimized and should not be used to estimate the runtime of your code on the hidden test set.</code></p>\n<p>I use kaggle notebook.</p>\n<p>if you need more information, please feel free to let me know.</p>",
      "rawMarkdown": "My model always predict that all score of the test data is 1. However my model seems to work on train data. Then I have no idea about how I can solve this problem.\n\n```python\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    df = pd.read_csv('path/to/my-dataset', index_col=[0])\n    \n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n    \n    env.predict(sample_submission)\n```\nIt might be a key to solve that this code raise an warning\n` This version of the API is not optimized and should not be used to estimate the runtime of your code on the hidden test set.`\n\nI use kaggle notebook.\n\nif you need more information, please feel free to let me know.",
      "votes": null
    },
    {
      "id": "2227983",
      "postDate": "04/20/2023 07:14:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/itsunori\" target=\"_blank\">@itsunori</a> , did you print p and best_threshold? maybe best_threshold is equal to zero? maybe p is always 1?</p>",
      "rawMarkdown": "Hi @itsunori , did you print p and best_threshold? maybe best_threshold is equal to zero? maybe p is always 1?",
      "votes": null
    },
    {
      "id": "2229500",
      "postDate": "04/21/2023 12:23:53",
      "content": "<p>What model do you use? </p>\n<p>Can you share more details? </p>\n<p>In general, it sounds like you might be experiencing overfitting, which means that while your model is performing well on your training data, it's not able to generalize well to new, unseen test data. </p>\n<p>Consider what's happening in that <code>for</code> loop. Without more information about the data and the models you're using, it's hard to say for sure, but it's possible that something in that loop is causing your model to always predict 1. You might try adding some print statements in there to see what's happening.</p>\n<p>I also noticed that you're using the <code>best_threshold</code> variable, which is not defined in the code you posted. Make sure that you've set this correctly – if <code>best_threshold</code> is to high or low it can make all predictions collaps.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "What model do you use? \n\nCan you share more details? \n\nIn general, it sounds like you might be experiencing overfitting, which means that while your model is performing well on your training data, it's not able to generalize well to new, unseen test data. \n\nConsider what's happening in that `for` loop. Without more information about the data and the models you're using, it's hard to say for sure, but it's possible that something in that loop is causing your model to always predict 1. You might try adding some print statements in there to see what's happening.\n\nI also noticed that you're using the `best_threshold` variable, which is not defined in the code you posted. Make sure that you've set this correctly – if `best_threshold` is to high or low it can make all predictions collaps.\n\nThe Devastator.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2227983,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "04/20/2023 07:14:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/itsunori\" target=\"_blank\">@itsunori</a> , did you print p and best_threshold? maybe best_threshold is equal to zero? maybe p is always 1?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2229500,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "04/21/2023 12:23:53",
      "content": "<p>What model do you use? </p>\n<p>Can you share more details? </p>\n<p>In general, it sounds like you might be experiencing overfitting, which means that while your model is performing well on your training data, it's not able to generalize well to new, unseen test data. </p>\n<p>Consider what's happening in that <code>for</code> loop. Without more information about the data and the models you're using, it's hard to say for sure, but it's possible that something in that loop is causing your model to always predict 1. You might try adding some print statements in there to see what's happening.</p>\n<p>I also noticed that you're using the <code>best_threshold</code> variable, which is not defined in the code you posted. Make sure that you've set this correctly – if <code>best_threshold</code> is to high or low it can make all predictions collaps.</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2226704": "My model always predict that all score of the test data is 1. However my model seems to work on train data. Then I have no idea about how I can solve this problem.\n\n```python\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    df = pd.read_csv('path/to/my-dataset', index_col=[0])\n    \n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n    \n    env.predict(sample_submission)\n```\nIt might be a key to solve that this code raise an warning\n` This version of the API is not optimized and should not be used to estimate the runtime of your code on the hidden test set.`\n\nI use kaggle notebook.\n\nif you need more information, please feel free to let me know.",
    "2227983": "Hi @itsunori , did you print p and best_threshold? maybe best_threshold is equal to zero? maybe p is always 1?",
    "2229500": "What model do you use? \n\nCan you share more details? \n\nIn general, it sounds like you might be experiencing overfitting, which means that while your model is performing well on your training data, it's not able to generalize well to new, unseen test data. \n\nConsider what's happening in that `for` loop. Without more information about the data and the models you're using, it's hard to say for sure, but it's possible that something in that loop is causing your model to always predict 1. You might try adding some print statements in there to see what's happening.\n\nI also noticed that you're using the `best_threshold` variable, which is not defined in the code you posted. Make sure that you've set this correctly – if `best_threshold` is to high or low it can make all predictions collaps.\n\nThe Devastator."
  },
  "source": "meta"
}