{
  "id": 408178,
  "title": "Submission error",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/408178",
  "author_name": "adriel cabral",
  "post_date": "2023-05-09T21:18:19.180000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I've been stuck for about 4 days in a 'Submission Score error' and it occurs when I add a 'new' feature in the 'feature_engineer_13_22' function that is also used in the other functions of the level_group 0-4 and 5-12, my notebook is not very different of this: <code>https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680</code></p>\n<p>The error does not occur when I use it only in level-groups 0-4 and 5-12.<br>\nCode of the new feature:</p>\n<pre><code>level_time = train.groupby([, ])[].agg().reset_index()\nlevel_time[] = level_time[].sub(level_time[].shift()).()\n\nlevel_time[] = level_time[] / \nlevel_time[] = level_time[] / \nlevel_time.loc[level_time[] == , ] = level_time.loc[level_time[] == , ]\n\nlevel_time_session_table = level_time.groupby([]).agg().reset_index().drop(columns=[, ])\n\nlevel_time_table = pd.pivot_table(level_time, values=[, ], index=, columns=[], fill_value=)\nlevel_time_table.columns = [.join((, col)).strip()  col  level_time_table.columns.values]\n\nlevel_time_table = level_time_table.reset_index().drop(columns=[])\n</code></pre>\n<p>Submission code</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n (test, sample_submission)  iter_test:\n\n    \n    \n    grp = test.level_group.values[]\n    a,b = limits[grp]\n\n     grp == :\n        df = feature_engineer_0_4(test.drop(columns=COLUMNS_DROP + []))\n     grp == :\n        df = feature_engineer_5_12(test.drop(columns=COLUMNS_DROP + []))\n     grp == :\n        df = feature_engineer_13_22(test.drop(columns=COLUMNS_DROP + []))\n        df = df.replace(np.nan, )\n\n     t  (a,b):\n        clf = models[]\n        p = clf.predict_proba(df.iloc[:, :].values.astype())[,]\n        mask = sample_submission.session_id..contains()\n        sample_submission.loc[mask,] = ( p &gt; best_threshold )\n        sample_submission = sample_submission.replace(np.nan, )\n\n    env.predict(sample_submission)\n</code></pre>\n<p>I tried to put zeros in possible NaNs (both in the submission and in the dataframe coming from the function), but it didn't solve the problem. I would appreciate it if anyone has any idea what could be causing it.</p>\n<p>Thank you for your attention ! </p>",
  "messages": [
    {
      "id": 2252065,
      "postDate": "2023-05-09T21:18:19.180Z",
      "content": "<p>I've been stuck for about 4 days in a 'Submission Score error' and it occurs when I add a 'new' feature in the 'feature_engineer_13_22' function that is also used in the other functions of the level_group 0-4 and 5-12, my notebook is not very different of this: <code>https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680</code></p>\n<p>The error does not occur when I use it only in level-groups 0-4 and 5-12.<br>\nCode of the new feature:</p>\n<pre><code>level_time = train.groupby([, ])[].agg().reset_index()\nlevel_time[] = level_time[].sub(level_time[].shift()).()\n\nlevel_time[] = level_time[] / \nlevel_time[] = level_time[] / \nlevel_time.loc[level_time[] == , ] = level_time.loc[level_time[] == , ]\n\nlevel_time_session_table = level_time.groupby([]).agg().reset_index().drop(columns=[, ])\n\nlevel_time_table = pd.pivot_table(level_time, values=[, ], index=, columns=[], fill_value=)\nlevel_time_table.columns = [.join((, col)).strip()  col  level_time_table.columns.values]\n\nlevel_time_table = level_time_table.reset_index().drop(columns=[])\n</code></pre>\n<p>Submission code</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n (test, sample_submission)  iter_test:\n\n    \n    \n    grp = test.level_group.values[]\n    a,b = limits[grp]\n\n     grp == :\n        df = feature_engineer_0_4(test.drop(columns=COLUMNS_DROP + []))\n     grp == :\n        df = feature_engineer_5_12(test.drop(columns=COLUMNS_DROP + []))\n     grp == :\n        df = feature_engineer_13_22(test.drop(columns=COLUMNS_DROP + []))\n        df = df.replace(np.nan, )\n\n     t  (a,b):\n        clf = models[]\n        p = clf.predict_proba(df.iloc[:, :].values.astype())[,]\n        mask = sample_submission.session_id..contains()\n        sample_submission.loc[mask,] = ( p &gt; best_threshold )\n        sample_submission = sample_submission.replace(np.nan, )\n\n    env.predict(sample_submission)\n</code></pre>\n<p>I tried to put zeros in possible NaNs (both in the submission and in the dataframe coming from the function), but it didn't solve the problem. I would appreciate it if anyone has any idea what could be causing it.</p>\n<p>Thank you for your attention ! </p>",
      "rawMarkdown": "I've been stuck for about 4 days in a 'Submission Score error' and it occurs when I add a 'new' feature in the 'feature_engineer_13_22' function that is also used in the other functions of the level_group 0-4 and 5-12, my notebook is not very different of this: `https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680`\n\nThe error does not occur when I use it only in level-groups 0-4 and 5-12.\nCode of the new feature:\n```python\n\nlevel_time = train.groupby(['session_id', 'level'])['elapsed_time'].agg('max').reset_index()\nlevel_time['diff'] = level_time['elapsed_time'].sub(level_time['elapsed_time'].shift(1)).abs()\n    \nlevel_time['elapsed_time'] = level_time['elapsed_time'] / 60000\nlevel_time['diff'] = level_time['diff'] / 60000\nlevel_time.loc[level_time['level'] == 13, 'diff'] = level_time.loc[level_time['level'] == 13, 'elapsed_time']\n    \nlevel_time_session_table = level_time.groupby(['session_id']).agg('median').reset_index().drop(columns=['level', 'session_id'])\n    \nlevel_time_table = pd.pivot_table(level_time, values=['elapsed_time', 'diff'], index='session_id', columns=['level'], fill_value=0)\nlevel_time_table.columns = ['_'.join(map(str, col)).strip() for col in level_time_table.columns.values]\n\nlevel_time_table = level_time_table.reset_index().drop(columns=['session_id'])\n```\nSubmission code\n\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nfor (test, sample_submission) in iter_test:\n    \n    # FEATURE ENGINEER TEST DATA\n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    \n    if grp == '0-4':\n        df = feature_engineer_0_4(test.drop(columns=COLUMNS_DROP + ['level_group']))\n    elif grp == '5-12':\n        df = feature_engineer_5_12(test.drop(columns=COLUMNS_DROP + ['level_group']))\n    elif grp == '13-22':\n        df = feature_engineer_13_22(test.drop(columns=COLUMNS_DROP + ['level_group']))\n        df = df.replace(np.nan, 0)\n    \n    for t in range(a,b):\n        clf = models[f'{t}']\n        p = clf.predict_proba(df.iloc[:, 1:].values.astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n        sample_submission = sample_submission.replace(np.nan, 0)\n    \n    env.predict(sample_submission)\n```\nI tried to put zeros in possible NaNs (both in the submission and in the dataframe coming from the function), but it didn't solve the problem. I would appreciate it if anyone has any idea what could be causing it.\n\nThank you for your attention ! "
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2252065": "I've been stuck for about 4 days in a 'Submission Score error' and it occurs when I add a 'new' feature in the 'feature_engineer_13_22' function that is also used in the other functions of the level_group 0-4 and 5-12, my notebook is not very different of this: `https://www.kaggle.com/code/cdeotte/xgboost-baseline-0-680`\n\nThe error does not occur when I use it only in level-groups 0-4 and 5-12.\nCode of the new feature:\n```python\n\nlevel_time = train.groupby(['session_id', 'level'])['elapsed_time'].agg('max').reset_index()\nlevel_time['diff'] = level_time['elapsed_time'].sub(level_time['elapsed_time'].shift(1)).abs()\n    \nlevel_time['elapsed_time'] = level_time['elapsed_time'] / 60000\nlevel_time['diff'] = level_time['diff'] / 60000\nlevel_time.loc[level_time['level'] == 13, 'diff'] = level_time.loc[level_time['level'] == 13, 'elapsed_time']\n    \nlevel_time_session_table = level_time.groupby(['session_id']).agg('median').reset_index().drop(columns=['level', 'session_id'])\n    \nlevel_time_table = pd.pivot_table(level_time, values=['elapsed_time', 'diff'], index='session_id', columns=['level'], fill_value=0)\nlevel_time_table.columns = ['_'.join(map(str, col)).strip() for col in level_time_table.columns.values]\n\nlevel_time_table = level_time_table.reset_index().drop(columns=['session_id'])\n```\nSubmission code\n\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nfor (test, sample_submission) in iter_test:\n    \n    # FEATURE ENGINEER TEST DATA\n    # INFER TEST DATA\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    \n    if grp == '0-4':\n        df = feature_engineer_0_4(test.drop(columns=COLUMNS_DROP + ['level_group']))\n    elif grp == '5-12':\n        df = feature_engineer_5_12(test.drop(columns=COLUMNS_DROP + ['level_group']))\n    elif grp == '13-22':\n        df = feature_engineer_13_22(test.drop(columns=COLUMNS_DROP + ['level_group']))\n        df = df.replace(np.nan, 0)\n    \n    for t in range(a,b):\n        clf = models[f'{t}']\n        p = clf.predict_proba(df.iloc[:, 1:].values.astype('float32'))[0,1]\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = int( p > best_threshold )\n        sample_submission = sample_submission.replace(np.nan, 0)\n    \n    env.predict(sample_submission)\n```\nI tried to put zeros in possible NaNs (both in the submission and in the dataframe coming from the function), but it didn't solve the problem. I would appreciate it if anyone has any idea what could be causing it.\n\nThank you for your attention ! "
  }
}