{
  "id": 396751,
  "title": "A way to submit without error",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/396751",
  "author_name": "",
  "post_date": "2023-03-22T18:45:33.972768300Z",
  "votes": 19,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I modified the prediction code as follows. Key is to explicitly use the question number equality when masking.</p>\n<pre><code> test, sample_submission  iter_test:\n    sample_submission[] = [(label.split()[][:])  label  sample_submission[]]\n...\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, ] = (p &gt; BEST_THRESHOLD).astype() \n...\n    env.predict(sample_submission[[, ]])\n</code></pre>",
  "messages": [
    {
      "id": "2192590",
      "postDate": "03/22/2023 18:45:33",
      "content": "<p>I modified the prediction code as follows. Key is to explicitly use the question number equality when masking.</p>\n<pre><code> test, sample_submission  iter_test:\n    sample_submission[] = [(label.split()[][:])  label  sample_submission[]]\n...\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, ] = (p &gt; BEST_THRESHOLD).astype() \n...\n    env.predict(sample_submission[[, ]])\n</code></pre>",
      "rawMarkdown": "I modified the prediction code as follows. Key is to explicitly use the question number equality when masking.\n\n```python\nfor test, sample_submission in iter_test:\n    sample_submission['question'] = [int(label.split('_')[1][1:]) for label in sample_submission['session_id']]\n...\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, 'correct'] = (p > BEST_THRESHOLD).astype('int') \n...\n    env.predict(sample_submission[['session_id', 'correct']])\n```",
      "votes": null
    },
    {
      "id": "2192623",
      "postDate": "03/22/2023 19:27:42",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, what is <code>t</code> in <code>mask = sample_submission.question == t</code>?</p>",
      "rawMarkdown": "cpmpml, what is `t` in `mask = sample_submission.question == t  `?",
      "votes": null
    },
    {
      "id": "2192672",
      "postDate": "03/22/2023 20:29:35",
      "content": "<p>the question number</p>",
      "rawMarkdown": "the question number",
      "votes": null
    },
    {
      "id": "2192673",
      "postDate": "03/22/2023 20:31:16",
      "content": "<p>for instance, in this notebook: <a href=\"https://www.kaggle.com/code/pourchot/simple-xgb\" target=\"_blank\">https://www.kaggle.com/code/pourchot/simple-xgb</a></p>",
      "rawMarkdown": "for instance, in this notebook: https://www.kaggle.com/code/pourchot/simple-xgb",
      "votes": null
    },
    {
      "id": "2200725",
      "postDate": "03/28/2023 18:45:32",
      "content": "<p>Yes had found and done the same. 😊 thanks for sharing but</p>",
      "rawMarkdown": "Yes had found and done the same. 😊 thanks for sharing but",
      "votes": null
    },
    {
      "id": "2201313",
      "postDate": "03/29/2023 08:08:10",
      "content": "<p>but what?       </p>",
      "rawMarkdown": "but what?",
      "votes": null
    },
    {
      "id": "2201671",
      "postDate": "03/29/2023 13:32:30",
      "content": "<p>do like this <a href=\"https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments#2195798\" target=\"_blank\">https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments#2195798</a> I mean </p>",
      "rawMarkdown": "do like this https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments#2195798 I mean",
      "votes": null
    },
    {
      "id": "2206479",
      "postDate": "04/02/2023 15:26:06",
      "content": "<p>Thanks for sharing, based on your answer I used this and it worked!</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n\n test, sample_submission  iter_test:\n    sample_submission[] = [(label.split()[][:])  label  sample_submission[]]\n    df = feature_engineer(test)\n\n    \n    grp = test.level_group.values[]\n    a,b = limits[grp]\n     t  (a,b):\n        clf = models[]\n        p = clf.predict_proba(df[FEATURES].astype())[:,]\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, ] = (p &gt; best_threshold).astype() \n    env.predict(sample_submission[[, ]])\n</code></pre>",
      "rawMarkdown": "Thanks for sharing, based on your answer I used this and it worked!\n\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor test, sample_submission in iter_test:\n    sample_submission['question'] = [int(label.split('_')[1][1:]) for label in sample_submission['session_id']]\n    df = feature_engineer(test)\n    \n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[:,1]\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, 'correct'] = (p > best_threshold).astype('int') \n    env.predict(sample_submission[['session_id', 'correct']])\n```",
      "votes": null
    },
    {
      "id": "2207416",
      "postDate": "04/03/2023 11:57:20",
      "content": "<p>very very Thank you~~!!</p>",
      "rawMarkdown": "very very Thank you~~!!",
      "votes": null
    },
    {
      "id": "2211098",
      "postDate": "04/05/2023 19:46:53",
      "content": "<p>It works perfectly, thanks a lot for sharing 🙌</p>",
      "rawMarkdown": "It works perfectly, thanks a lot for sharing 🙌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2192623,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/22/2023 19:27:42",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, what is <code>t</code> in <code>mask = sample_submission.question == t</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2192672,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/22/2023 20:29:35",
          "content": "<p>the question number</p>",
          "votes": null,
          "replies": [
            {
              "id": 2192673,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "03/22/2023 20:31:16",
              "content": "<p>for instance, in this notebook: <a href=\"https://www.kaggle.com/code/pourchot/simple-xgb\" target=\"_blank\">https://www.kaggle.com/code/pourchot/simple-xgb</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2200725,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "03/28/2023 18:45:32",
      "content": "<p>Yes had found and done the same. 😊 thanks for sharing but</p>",
      "votes": null,
      "replies": [
        {
          "id": 2201313,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/29/2023 08:08:10",
          "content": "<p>but what?       </p>",
          "votes": null,
          "replies": [
            {
              "id": 2201671,
              "author_name": "gauravbrills",
              "author_url": "",
              "post_date": "03/29/2023 13:32:30",
              "content": "<p>do like this <a href=\"https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments#2195798\" target=\"_blank\">https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments#2195798</a> I mean </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2206479,
      "author_name": "jpardom",
      "author_url": "",
      "post_date": "04/02/2023 15:26:06",
      "content": "<p>Thanks for sharing, based on your answer I used this and it worked!</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n\n test, sample_submission  iter_test:\n    sample_submission[] = [(label.split()[][:])  label  sample_submission[]]\n    df = feature_engineer(test)\n\n    \n    grp = test.level_group.values[]\n    a,b = limits[grp]\n     t  (a,b):\n        clf = models[]\n        p = clf.predict_proba(df[FEATURES].astype())[:,]\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, ] = (p &gt; best_threshold).astype() \n    env.predict(sample_submission[[, ]])\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2211098,
          "author_name": "javihm77",
          "author_url": "",
          "post_date": "04/05/2023 19:46:53",
          "content": "<p>It works perfectly, thanks a lot for sharing 🙌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2207416,
      "author_name": "bsj940528",
      "author_url": "",
      "post_date": "04/03/2023 11:57:20",
      "content": "<p>very very Thank you~~!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2192590": "I modified the prediction code as follows. Key is to explicitly use the question number equality when masking.\n\n```python\nfor test, sample_submission in iter_test:\n    sample_submission['question'] = [int(label.split('_')[1][1:]) for label in sample_submission['session_id']]\n...\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, 'correct'] = (p > BEST_THRESHOLD).astype('int') \n...\n    env.predict(sample_submission[['session_id', 'correct']])\n```",
    "2192623": "cpmpml, what is `t` in `mask = sample_submission.question == t  `?",
    "2192672": "the question number",
    "2192673": "for instance, in this notebook: https://www.kaggle.com/code/pourchot/simple-xgb",
    "2200725": "Yes had found and done the same. 😊 thanks for sharing but",
    "2201313": "but what?",
    "2201671": "do like this https://www.kaggle.com/code/leehomhuang/catboost-baseline-with-lots-features-inference/comments#2195798 I mean",
    "2206479": "Thanks for sharing, based on your answer I used this and it worked!\n\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor test, sample_submission in iter_test:\n    sample_submission['question'] = [int(label.split('_')[1][1:]) for label in sample_submission['session_id']]\n    df = feature_engineer(test)\n    \n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        clf = models[f'{grp}_{t}']\n        p = clf.predict_proba(df[FEATURES].astype('float32'))[:,1]\n        mask = sample_submission.question == t    \n        sample_submission.loc[mask, 'correct'] = (p > best_threshold).astype('int') \n    env.predict(sample_submission[['session_id', 'correct']])\n```",
    "2207416": "very very Thank you~~!!",
    "2211098": "It works perfectly, thanks a lot for sharing 🙌"
  },
  "source": "meta"
}