{
  "id": 402887,
  "title": "19th place solution, stacking 3 models",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402887",
  "author_name": "1110Ra",
  "post_date": "2023-04-20T04:59:39.904000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Greetings, </p>\n<p>In this Kaggle competition, I divided the dataset into four segments and trained four different models. Each part was then evaluated using the other three models, and the prediction error was measured against the target values. The distribution of errors for all events is shown below, revealing distinct peaks for both signal and noise events.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8459123%2Ff88c90293b1794b0e4b592eb775d6c8f%2Ferror_dist.png?generation=1681964940582146&amp;alt=media\" alt=\"\"></p>\n<p>To proceed, I randomly selected from signal events to create two separate databases and generated one database for noise events. <br>\nI modified the database generator code to preserve the initial 600 pulses of every event. You can find the updated <strong>notebook</strong> <a href=\"https://www.kaggle.com/mohammadrahmati/efficient-database-generator\" target=\"_blank\">here</a>. Here are the main differences compared to the original code:</p>\n<ul>\n<li>The revised code generates more compact databases, with the three databases totaling 809 GB instead of the original 1.5 TB database, despite having three times more pulses. This reduction in size is achieved by eliminating extraneous data.</li>\n<li>The revised code generates database based on the list of event_ids and batch_ids, </li>\n<li>retaining all events while saving the first 600 pulses of each, </li>\n<li>producing train and validation event_id sets, which are then saved in pickle format.</li>\n</ul>\n<p>The models were trained, resulting in the following training error, validation error, and LB score:</p>\n<ul>\n<li>Signal Part 1: 49 epochs, training error: 0.599, validation error: 0.616, LB_score: 0.995</li>\n<li>Signal Part 2: 60 epochs, training error: 0.622, validation error: 0.644, LB_score: 0.997</li>\n<li>Noise: 40 epochs, training error: 2.43, validation error: 2.43</li>\n</ul>\n<p>Finally, I combined the models using the following stacking logic:</p>\n<pre><code> ():\n    df = pd.concat(dfs)\n    df[] = df[] * df[]\n    df[] = df[] * df[]\n    df[] = df[] * df[]\n\n    summed = df.groupby().agg({: , : , : , : })\n    summed[] = summed[] / summed[]\n    summed[] = summed[] / summed[]\n    summed[] = summed[] / summed[]\n    summed.rename({: , : , : }, axis=, inplace=)\n    summed.reset_index(inplace=)\n\n     summed[[, , , , ]]\n</code></pre>\n<p><code>result_agg = sum_kappa_times_xyz(pred_signal_p1, pred_signal_p2)</code></p>\n<pre><code>threshold = \npred_noise.loc[result_agg.direction_kappa &gt; threshold, ] = \nresult_agg = sum_kappa_times_xyz(pred_noise, result_agg)\n</code></pre>\n<pre><code>submission_df = prepare_dataframe(result_agg, angle_post_fix = )\nsubmission_df.to_csv()\n</code></pre>",
  "messages": [
    {
      "id": 2227883,
      "postDate": "2023-04-20T04:59:39.903Z",
      "content": "<p>Greetings, </p>\n<p>In this Kaggle competition, I divided the dataset into four segments and trained four different models. Each part was then evaluated using the other three models, and the prediction error was measured against the target values. The distribution of errors for all events is shown below, revealing distinct peaks for both signal and noise events.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8459123%2Ff88c90293b1794b0e4b592eb775d6c8f%2Ferror_dist.png?generation=1681964940582146&amp;alt=media\" alt=\"\"></p>\n<p>To proceed, I randomly selected from signal events to create two separate databases and generated one database for noise events. <br>\nI modified the database generator code to preserve the initial 600 pulses of every event. You can find the updated <strong>notebook</strong> <a href=\"https://www.kaggle.com/mohammadrahmati/efficient-database-generator\" target=\"_blank\">here</a>. Here are the main differences compared to the original code:</p>\n<ul>\n<li>The revised code generates more compact databases, with the three databases totaling 809 GB instead of the original 1.5 TB database, despite having three times more pulses. This reduction in size is achieved by eliminating extraneous data.</li>\n<li>The revised code generates database based on the list of event_ids and batch_ids, </li>\n<li>retaining all events while saving the first 600 pulses of each, </li>\n<li>producing train and validation event_id sets, which are then saved in pickle format.</li>\n</ul>\n<p>The models were trained, resulting in the following training error, validation error, and LB score:</p>\n<ul>\n<li>Signal Part 1: 49 epochs, training error: 0.599, validation error: 0.616, LB_score: 0.995</li>\n<li>Signal Part 2: 60 epochs, training error: 0.622, validation error: 0.644, LB_score: 0.997</li>\n<li>Noise: 40 epochs, training error: 2.43, validation error: 2.43</li>\n</ul>\n<p>Finally, I combined the models using the following stacking logic:</p>\n<pre><code> ():\n    df = pd.concat(dfs)\n    df[] = df[] * df[]\n    df[] = df[] * df[]\n    df[] = df[] * df[]\n\n    summed = df.groupby().agg({: , : , : , : })\n    summed[] = summed[] / summed[]\n    summed[] = summed[] / summed[]\n    summed[] = summed[] / summed[]\n    summed.rename({: , : , : }, axis=, inplace=)\n    summed.reset_index(inplace=)\n\n     summed[[, , , , ]]\n</code></pre>\n<p><code>result_agg = sum_kappa_times_xyz(pred_signal_p1, pred_signal_p2)</code></p>\n<pre><code>threshold = \npred_noise.loc[result_agg.direction_kappa &gt; threshold, ] = \nresult_agg = sum_kappa_times_xyz(pred_noise, result_agg)\n</code></pre>\n<pre><code>submission_df = prepare_dataframe(result_agg, angle_post_fix = )\nsubmission_df.to_csv()\n</code></pre>",
      "rawMarkdown": "Greetings, \n\nIn this Kaggle competition, I divided the dataset into four segments and trained four different models. Each part was then evaluated using the other three models, and the prediction error was measured against the target values. The distribution of errors for all events is shown below, revealing distinct peaks for both signal and noise events.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8459123%2Ff88c90293b1794b0e4b592eb775d6c8f%2Ferror_dist.png?generation=1681964940582146&alt=media)\n\nTo proceed, I randomly selected from signal events to create two separate databases and generated one database for noise events. \nI modified the database generator code to preserve the initial 600 pulses of every event. You can find the updated **notebook** [here](https://www.kaggle.com/mohammadrahmati/efficient-database-generator). Here are the main differences compared to the original code:\n- The revised code generates more compact databases, with the three databases totaling 809 GB instead of the original 1.5 TB database, despite having three times more pulses. This reduction in size is achieved by eliminating extraneous data.\n- The revised code generates database based on the list of event_ids and batch_ids, \n- retaining all events while saving the first 600 pulses of each, \n- producing train and validation event_id sets, which are then saved in pickle format.\n\nThe models were trained, resulting in the following training error, validation error, and LB score:\n- Signal Part 1: 49 epochs, training error: 0.599, validation error: 0.616, LB_score: 0.995\n- Signal Part 2: 60 epochs, training error: 0.622, validation error: 0.644, LB_score: 0.997\n- Noise: 40 epochs, training error: 2.43, validation error: 2.43\n\nFinally, I combined the models using the following stacking logic:\n\n```python\ndef sum_kappa_times_xyz(*dfs):\n    df = pd.concat(dfs)\n    df['kappa_times_x'] = df['direction_kappa'] * df['direction_x']\n    df['kappa_times_y'] = df['direction_kappa'] * df['direction_y']\n    df['kappa_times_z'] = df['direction_kappa'] * df['direction_z']\n    \n    summed = df.groupby('event_id').agg({'kappa_times_x': 'sum', 'kappa_times_y': 'sum', 'kappa_times_z': 'sum', 'direction_kappa': 'sum'})\n    summed['kappa_times_x'] = summed['kappa_times_x'] / summed['direction_kappa']\n    summed['kappa_times_y'] = summed['kappa_times_y'] / summed['direction_kappa']\n    summed['kappa_times_z'] = summed['kappa_times_z'] / summed['direction_kappa']\n    summed.rename({'kappa_times_x': 'direction_x', 'kappa_times_y': 'direction_y', 'kappa_times_z': 'direction_z'}, axis=1, inplace=True)\n    summed.reset_index(inplace=True)\n    \n    return summed[['direction_x', 'direction_y', 'direction_z', 'direction_kappa', 'event_id']]\n```\n\n`result_agg = sum_kappa_times_xyz(pred_signal_p1, pred_signal_p2)`\n\n```python\nthreshold = 10\npred_noise.loc[result_agg.direction_kappa > threshold, 'direction_kappa'] = 0\nresult_agg = sum_kappa_times_xyz(pred_noise, result_agg)\n```\n\n```python\nsubmission_df = prepare_dataframe(result_agg, angle_post_fix = '')\nsubmission_df.to_csv('submission.csv')\n``` ",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2227883": "Greetings, \n\nIn this Kaggle competition, I divided the dataset into four segments and trained four different models. Each part was then evaluated using the other three models, and the prediction error was measured against the target values. The distribution of errors for all events is shown below, revealing distinct peaks for both signal and noise events.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8459123%2Ff88c90293b1794b0e4b592eb775d6c8f%2Ferror_dist.png?generation=1681964940582146&alt=media)\n\nTo proceed, I randomly selected from signal events to create two separate databases and generated one database for noise events. \nI modified the database generator code to preserve the initial 600 pulses of every event. You can find the updated **notebook** [here](https://www.kaggle.com/mohammadrahmati/efficient-database-generator). Here are the main differences compared to the original code:\n- The revised code generates more compact databases, with the three databases totaling 809 GB instead of the original 1.5 TB database, despite having three times more pulses. This reduction in size is achieved by eliminating extraneous data.\n- The revised code generates database based on the list of event_ids and batch_ids, \n- retaining all events while saving the first 600 pulses of each, \n- producing train and validation event_id sets, which are then saved in pickle format.\n\nThe models were trained, resulting in the following training error, validation error, and LB score:\n- Signal Part 1: 49 epochs, training error: 0.599, validation error: 0.616, LB_score: 0.995\n- Signal Part 2: 60 epochs, training error: 0.622, validation error: 0.644, LB_score: 0.997\n- Noise: 40 epochs, training error: 2.43, validation error: 2.43\n\nFinally, I combined the models using the following stacking logic:\n\n```python\ndef sum_kappa_times_xyz(*dfs):\n    df = pd.concat(dfs)\n    df['kappa_times_x'] = df['direction_kappa'] * df['direction_x']\n    df['kappa_times_y'] = df['direction_kappa'] * df['direction_y']\n    df['kappa_times_z'] = df['direction_kappa'] * df['direction_z']\n    \n    summed = df.groupby('event_id').agg({'kappa_times_x': 'sum', 'kappa_times_y': 'sum', 'kappa_times_z': 'sum', 'direction_kappa': 'sum'})\n    summed['kappa_times_x'] = summed['kappa_times_x'] / summed['direction_kappa']\n    summed['kappa_times_y'] = summed['kappa_times_y'] / summed['direction_kappa']\n    summed['kappa_times_z'] = summed['kappa_times_z'] / summed['direction_kappa']\n    summed.rename({'kappa_times_x': 'direction_x', 'kappa_times_y': 'direction_y', 'kappa_times_z': 'direction_z'}, axis=1, inplace=True)\n    summed.reset_index(inplace=True)\n    \n    return summed[['direction_x', 'direction_y', 'direction_z', 'direction_kappa', 'event_id']]\n```\n\n`result_agg = sum_kappa_times_xyz(pred_signal_p1, pred_signal_p2)`\n\n```python\nthreshold = 10\npred_noise.loc[result_agg.direction_kappa > threshold, 'direction_kappa'] = 0\nresult_agg = sum_kappa_times_xyz(pred_noise, result_agg)\n```\n\n```python\nsubmission_df = prepare_dataframe(result_agg, angle_post_fix = '')\nsubmission_df.to_csv('submission.csv')\n``` "
  }
}