{
  "id": 552625,
  "title": "Private 7th place solution",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552625",
  "author_name": "sqrt4kaido",
  "post_date": "2024-12-20T16:24:11.799000",
  "votes": 31,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello, everyone. </p>\n<p>First, I want to thank the organizers for hosting this competition. Working with real-world data was challenging but provided an excellent opportunity to enhance my technical skills for practical applications.</p>\n<p>Below, I share my solution:</p>\n<h3>Model</h3>\n<ul>\n<li>I used this baseline: <a href=\"https://www.kaggle.com/code/greysky/cmi-single-lgbm-cv-0-471-lb-0-460\" target=\"_blank\">CMI Single LGBM CV 0.471 LB 0.460</a>.</li>\n<li>The primary model is LGBM(, with XGBoost included in the ensemble).</li>\n<li>Missing values were handled using median imputation, calculated only on the training data for each fold, and applied to the validation/test sets.</li>\n<li>Sequential data statistics were processed as in the public notebook but optimized for speed using Polars.</li>\n<li>I used Tweedie loss as the primary loss function(, while classification was also included in the ensemble). → <strong>important</strong></li>\n<li>Predictions for data where 'sii' was not present were generated and used as pseudo-labels for training. → <strong>important</strong></li>\n<li>Features defined as categorical integers in the data dictionary were used both categorically and numerically.</li>\n</ul>\n<h3>Threshold Optimization → <strong>important</strong></h3>\n<p>Threshold optimization was performed on each fold's training data and applied to the validation/test sets. Using percentiles further improved both accuracy and robustness. Initial optimization values were derived from the discussion <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551533\" target=\"_blank\">here</a>.</p>\n<pre><code> ():\n    tmp_df = pd.DataFrame({: y, : y_pred})\n    oof_initial_thresholds = (\n        tmp_df.groupby()[].mean().iloc[:].values.tolist()\n    )\n\n    oof_threshold_percentiles = [\n        (tmp_df[] &lt;= threshold).mean()\n         threshold  oof_initial_thresholds\n    ]\n\n     oof_threshold_percentiles\n\noof_initial_thresholds = calc_initial_th(y, y_pred)\n\n.optimizer = minimize(\n    eval_preds_percentile,\n    x0=oof_initial_thresholds,\n    args=(y, y_pred),\n    method=,\n    bounds=[(, ), (, ), (, )],\n)\n</code></pre>\n<h3>CV Strategy and Model Evaluation</h3>\n<p>I used 5-fold StratifiedKFold. However, as you may know, scores can vary significantly depending on the seed. In the aforementioned notebook, the results were averaged across 10 seeds for submission. Similarly, in validation, I evaluated the voting results from 10 seeds as the score for each experiment. I also evaluated the 10-seed results in validation and optimized parameters for each seed using Optuna.</p>\n<p>For example, one submission produced the following seed-dependent score variations:</p>\n<pre><code>\n</code></pre>\n<p>I monitored both the voting score and the average of the individual scores. For instance, in the example above, the voting score is 0.49218340134866767, and the average score is 0.4839108063626429. For submissions, I performed a voting ensemble using 5 folds × 10 seeds per model.</p>\n<p>While I aimed to make the evaluation and learning process as robust as possible, I think my ranking still depended heavily on luck.<br>\nThank you for reading!</p>",
  "messages": [
    {
      "id": 3077162,
      "postDate": "2024-12-20T16:24:11.800Z",
      "content": "<p>Hello, everyone. </p>\n<p>First, I want to thank the organizers for hosting this competition. Working with real-world data was challenging but provided an excellent opportunity to enhance my technical skills for practical applications.</p>\n<p>Below, I share my solution:</p>\n<h3>Model</h3>\n<ul>\n<li>I used this baseline: <a href=\"https://www.kaggle.com/code/greysky/cmi-single-lgbm-cv-0-471-lb-0-460\" target=\"_blank\">CMI Single LGBM CV 0.471 LB 0.460</a>.</li>\n<li>The primary model is LGBM(, with XGBoost included in the ensemble).</li>\n<li>Missing values were handled using median imputation, calculated only on the training data for each fold, and applied to the validation/test sets.</li>\n<li>Sequential data statistics were processed as in the public notebook but optimized for speed using Polars.</li>\n<li>I used Tweedie loss as the primary loss function(, while classification was also included in the ensemble). → <strong>important</strong></li>\n<li>Predictions for data where 'sii' was not present were generated and used as pseudo-labels for training. → <strong>important</strong></li>\n<li>Features defined as categorical integers in the data dictionary were used both categorically and numerically.</li>\n</ul>\n<h3>Threshold Optimization → <strong>important</strong></h3>\n<p>Threshold optimization was performed on each fold's training data and applied to the validation/test sets. Using percentiles further improved both accuracy and robustness. Initial optimization values were derived from the discussion <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551533\" target=\"_blank\">here</a>.</p>\n<pre><code> ():\n    tmp_df = pd.DataFrame({: y, : y_pred})\n    oof_initial_thresholds = (\n        tmp_df.groupby()[].mean().iloc[:].values.tolist()\n    )\n\n    oof_threshold_percentiles = [\n        (tmp_df[] &lt;= threshold).mean()\n         threshold  oof_initial_thresholds\n    ]\n\n     oof_threshold_percentiles\n\noof_initial_thresholds = calc_initial_th(y, y_pred)\n\n.optimizer = minimize(\n    eval_preds_percentile,\n    x0=oof_initial_thresholds,\n    args=(y, y_pred),\n    method=,\n    bounds=[(, ), (, ), (, )],\n)\n</code></pre>\n<h3>CV Strategy and Model Evaluation</h3>\n<p>I used 5-fold StratifiedKFold. However, as you may know, scores can vary significantly depending on the seed. In the aforementioned notebook, the results were averaged across 10 seeds for submission. Similarly, in validation, I evaluated the voting results from 10 seeds as the score for each experiment. I also evaluated the 10-seed results in validation and optimized parameters for each seed using Optuna.</p>\n<p>For example, one submission produced the following seed-dependent score variations:</p>\n<pre><code>\n</code></pre>\n<p>I monitored both the voting score and the average of the individual scores. For instance, in the example above, the voting score is 0.49218340134866767, and the average score is 0.4839108063626429. For submissions, I performed a voting ensemble using 5 folds × 10 seeds per model.</p>\n<p>While I aimed to make the evaluation and learning process as robust as possible, I think my ranking still depended heavily on luck.<br>\nThank you for reading!</p>",
      "rawMarkdown": "Hello, everyone. \n\nFirst, I want to thank the organizers for hosting this competition. Working with real-world data was challenging but provided an excellent opportunity to enhance my technical skills for practical applications.\n\nBelow, I share my solution:\n\n### Model\n\n- I used this baseline: [CMI Single LGBM CV 0.471 LB 0.460](https://www.kaggle.com/code/greysky/cmi-single-lgbm-cv-0-471-lb-0-460).\n- The primary model is LGBM(, with XGBoost included in the ensemble).\n- Missing values were handled using median imputation, calculated only on the training data for each fold, and applied to the validation/test sets.\n- Sequential data statistics were processed as in the public notebook but optimized for speed using Polars.\n- I used Tweedie loss as the primary loss function(, while classification was also included in the ensemble). → **important**\n- Predictions for data where 'sii' was not present were generated and used as pseudo-labels for training. → **important**\n- Features defined as categorical integers in the data dictionary were used both categorically and numerically.\n\n### Threshold Optimization → **important**\n\nThreshold optimization was performed on each fold's training data and applied to the validation/test sets. Using percentiles further improved both accuracy and robustness. Initial optimization values were derived from the discussion [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551533).\n\n```python\ndef calc_initial_th(y, y_pred):\n    tmp_df = pd.DataFrame({\"sii\": y, \"prediction\": y_pred})\n    oof_initial_thresholds = (\n        tmp_df.groupby(\"sii\")[\"prediction\"].mean().iloc[1:].values.tolist()\n    )\n\n    oof_threshold_percentiles = [\n        (tmp_df[\"prediction\"] <= threshold).mean()\n        for threshold in oof_initial_thresholds\n    ]\n\n    return oof_threshold_percentiles\n\noof_initial_thresholds = calc_initial_th(y, y_pred)\n\nself.optimizer = minimize(\n    eval_preds_percentile,\n    x0=oof_initial_thresholds,\n    args=(y, y_pred),\n    method=\"Nelder-Mead\",\n    bounds=[(0, 1), (0, 1), (0, 1)],\n)\n```\n\n### CV Strategy and Model Evaluation\n\nI used 5-fold StratifiedKFold. However, as you may know, scores can vary significantly depending on the seed. In the aforementioned notebook, the results were averaged across 10 seeds for submission. Similarly, in validation, I evaluated the voting results from 10 seeds as the score for each experiment. I also evaluated the 10-seed results in validation and optimized parameters for each seed using Optuna.\n\nFor example, one submission produced the following seed-dependent score variations:\n```\n[0.4924600384356978,\n 0.480933892862644,\n 0.4896986007283455,\n 0.4848801590924426,\n 0.47931930137389445,\n 0.48348085393094636,\n 0.48327061771577756,\n 0.47943771383186834,\n 0.47465614619853563,\n 0.4909707394562764]\n```\nI monitored both the voting score and the average of the individual scores. For instance, in the example above, the voting score is 0.49218340134866767, and the average score is 0.4839108063626429. For submissions, I performed a voting ensemble using 5 folds × 10 seeds per model.\n\n\nWhile I aimed to make the evaluation and learning process as robust as possible, I think my ranking still depended heavily on luck.\nThank you for reading!\n",
      "votes": 31
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3077162": "Hello, everyone. \n\nFirst, I want to thank the organizers for hosting this competition. Working with real-world data was challenging but provided an excellent opportunity to enhance my technical skills for practical applications.\n\nBelow, I share my solution:\n\n### Model\n\n- I used this baseline: [CMI Single LGBM CV 0.471 LB 0.460](https://www.kaggle.com/code/greysky/cmi-single-lgbm-cv-0-471-lb-0-460).\n- The primary model is LGBM(, with XGBoost included in the ensemble).\n- Missing values were handled using median imputation, calculated only on the training data for each fold, and applied to the validation/test sets.\n- Sequential data statistics were processed as in the public notebook but optimized for speed using Polars.\n- I used Tweedie loss as the primary loss function(, while classification was also included in the ensemble). → **important**\n- Predictions for data where 'sii' was not present were generated and used as pseudo-labels for training. → **important**\n- Features defined as categorical integers in the data dictionary were used both categorically and numerically.\n\n### Threshold Optimization → **important**\n\nThreshold optimization was performed on each fold's training data and applied to the validation/test sets. Using percentiles further improved both accuracy and robustness. Initial optimization values were derived from the discussion [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551533).\n\n```python\ndef calc_initial_th(y, y_pred):\n    tmp_df = pd.DataFrame({\"sii\": y, \"prediction\": y_pred})\n    oof_initial_thresholds = (\n        tmp_df.groupby(\"sii\")[\"prediction\"].mean().iloc[1:].values.tolist()\n    )\n\n    oof_threshold_percentiles = [\n        (tmp_df[\"prediction\"] <= threshold).mean()\n        for threshold in oof_initial_thresholds\n    ]\n\n    return oof_threshold_percentiles\n\noof_initial_thresholds = calc_initial_th(y, y_pred)\n\nself.optimizer = minimize(\n    eval_preds_percentile,\n    x0=oof_initial_thresholds,\n    args=(y, y_pred),\n    method=\"Nelder-Mead\",\n    bounds=[(0, 1), (0, 1), (0, 1)],\n)\n```\n\n### CV Strategy and Model Evaluation\n\nI used 5-fold StratifiedKFold. However, as you may know, scores can vary significantly depending on the seed. In the aforementioned notebook, the results were averaged across 10 seeds for submission. Similarly, in validation, I evaluated the voting results from 10 seeds as the score for each experiment. I also evaluated the 10-seed results in validation and optimized parameters for each seed using Optuna.\n\nFor example, one submission produced the following seed-dependent score variations:\n```\n[0.4924600384356978,\n 0.480933892862644,\n 0.4896986007283455,\n 0.4848801590924426,\n 0.47931930137389445,\n 0.48348085393094636,\n 0.48327061771577756,\n 0.47943771383186834,\n 0.47465614619853563,\n 0.4909707394562764]\n```\nI monitored both the voting score and the average of the individual scores. For instance, in the example above, the voting score is 0.49218340134866767, and the average score is 0.4839108063626429. For submissions, I performed a voting ensemble using 5 folds × 10 seeds per model.\n\n\nWhile I aimed to make the evaluation and learning process as robust as possible, I think my ranking still depended heavily on luck.\nThank you for reading!\n"
  }
}