{
  "id": 553146,
  "title": "6th place solution",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/553146",
  "author_name": "m_furu",
  "post_date": "2024-12-24T01:34:26.911000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello to all the competition participants. <br>\nThank you to the competition organizers for providing us with this valuable experience.</p>\n<p>I gave up trying to catch up with the more skilled participants two months ago.  I am very surprised and confused by the change in ranking after I gave up.  I recognize this result is due to luck, not my ability.</p>\n<p>I will share the solution that gave me such unexpected results below.</p>\n<p><strong>Overview</strong></p>\n<ul>\n<li>Models<br>\nThree models ensemble (simple average).<br>\nThe models are Vision Transformer (change the input layers code), lightgbm, catboost.</li>\n<li>Preprocessing<br>\nAggregate parquet data into one row per id using mean, std, etc.<br>\nConvert categorical variables using 'to_dummies (polars)'. <br>\nImpute null values ​​using 'group_by('Basic_Demos-Age', 'Basic_Demos-Sex').mean()' in the training data.<br>\nStandardize features using min_max of the training data.</li>\n<li>Learning<br>\nMetric : mae<br>\nPerforme cross-validation using 'StratifiedKFold (n_splits=5, y:'sii')'.<br>\nLike many other participants, Use 'threshold_rounder' (after the ensemble).</li>\n<li>Notebook<br>\n<a href=\"https://www.kaggle.com/code/miyafuru/internet-use-v2?scriptVersionId=201867541\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>Result</strong><br>\n　CV(before threshold_rounder) : 0.408,  CV(after threshold_rounder) : 0.481,  Public LB : 0.471,  Private LB : 0.476</p>\n<p><strong>Changes that worked</strong></p>\n<ul>\n<li>Using the Transformer<br>\nBest result before use（twe models ensemble）<br>\nCV(before threshold_rounder) : 0.389,  CV(after threshold_rounder) : 0.483,  Public LB : 0.463,  Private LB : 0.472</li>\n</ul>\n<p><strong>Changes that didn't work</strong></p>\n<ul>\n<li>Optimizing the ensemble weights<br>\nBest result<br>\nCV(before threshold_rounder) : 0.398,  CV(after threshold_rounder) : 0.483,  Public LB : 0.454,  Private LB : 0.463</li>\n<li>Metric : quadratic weighted kappa<br>\nBest result<br>\nCV(before threshold_rounder) : 0.429,  CV(after threshold_rounder) : 0.485,  Public LB : 0.457,  Private LB : 0.470</li>\n<li>Using 'PCIAT-PCIAT_Total' as target<br>\nBest result<br>\nCV(before threshold_rounder) : not calculated,  CV(after threshold_rounder) : 0.481,  Public LB : 0.461,  Private LB : 0.471</li>\n</ul>\n<p>I hope this post helps you understand the surprising ranking changes.<br>\nThank you for reading !</p>",
  "messages": [
    {
      "id": 3079674,
      "postDate": "2024-12-24T01:34:26.910Z",
      "content": "<p>Hello to all the competition participants. <br>\nThank you to the competition organizers for providing us with this valuable experience.</p>\n<p>I gave up trying to catch up with the more skilled participants two months ago.  I am very surprised and confused by the change in ranking after I gave up.  I recognize this result is due to luck, not my ability.</p>\n<p>I will share the solution that gave me such unexpected results below.</p>\n<p><strong>Overview</strong></p>\n<ul>\n<li>Models<br>\nThree models ensemble (simple average).<br>\nThe models are Vision Transformer (change the input layers code), lightgbm, catboost.</li>\n<li>Preprocessing<br>\nAggregate parquet data into one row per id using mean, std, etc.<br>\nConvert categorical variables using 'to_dummies (polars)'. <br>\nImpute null values ​​using 'group_by('Basic_Demos-Age', 'Basic_Demos-Sex').mean()' in the training data.<br>\nStandardize features using min_max of the training data.</li>\n<li>Learning<br>\nMetric : mae<br>\nPerforme cross-validation using 'StratifiedKFold (n_splits=5, y:'sii')'.<br>\nLike many other participants, Use 'threshold_rounder' (after the ensemble).</li>\n<li>Notebook<br>\n<a href=\"https://www.kaggle.com/code/miyafuru/internet-use-v2?scriptVersionId=201867541\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>Result</strong><br>\n　CV(before threshold_rounder) : 0.408,  CV(after threshold_rounder) : 0.481,  Public LB : 0.471,  Private LB : 0.476</p>\n<p><strong>Changes that worked</strong></p>\n<ul>\n<li>Using the Transformer<br>\nBest result before use（twe models ensemble）<br>\nCV(before threshold_rounder) : 0.389,  CV(after threshold_rounder) : 0.483,  Public LB : 0.463,  Private LB : 0.472</li>\n</ul>\n<p><strong>Changes that didn't work</strong></p>\n<ul>\n<li>Optimizing the ensemble weights<br>\nBest result<br>\nCV(before threshold_rounder) : 0.398,  CV(after threshold_rounder) : 0.483,  Public LB : 0.454,  Private LB : 0.463</li>\n<li>Metric : quadratic weighted kappa<br>\nBest result<br>\nCV(before threshold_rounder) : 0.429,  CV(after threshold_rounder) : 0.485,  Public LB : 0.457,  Private LB : 0.470</li>\n<li>Using 'PCIAT-PCIAT_Total' as target<br>\nBest result<br>\nCV(before threshold_rounder) : not calculated,  CV(after threshold_rounder) : 0.481,  Public LB : 0.461,  Private LB : 0.471</li>\n</ul>\n<p>I hope this post helps you understand the surprising ranking changes.<br>\nThank you for reading !</p>",
      "rawMarkdown": "Hello to all the competition participants. \nThank you to the competition organizers for providing us with this valuable experience.\n\nI gave up trying to catch up with the more skilled participants two months ago.  I am very surprised and confused by the change in ranking after I gave up.  I recognize this result is due to luck, not my ability.\n\nI will share the solution that gave me such unexpected results below.\n\n**Overview**\n- Models\nThree models ensemble (simple average).\nThe models are Vision Transformer (change the input layers code), lightgbm, catboost.\n- Preprocessing\nAggregate parquet data into one row per id using mean, std, etc.\nConvert categorical variables using 'to_dummies (polars)'. \nImpute null values ​​using 'group_by('Basic_Demos-Age', 'Basic_Demos-Sex').mean()' in the training data.\nStandardize features using min_max of the training data.\n- Learning\nMetric : mae\nPerforme cross-validation using 'StratifiedKFold (n_splits=5, y:'sii')'.\nLike many other participants, Use 'threshold_rounder' (after the ensemble).\n- Notebook\n[here](https://www.kaggle.com/code/miyafuru/internet-use-v2?scriptVersionId=201867541)\n\n**Result**\n　CV(before threshold_rounder) : 0.408,  CV(after threshold_rounder) : 0.481,  Public LB : 0.471,  Private LB : 0.476\n\n**Changes that worked**\n- Using the Transformer\nBest result before use（twe models ensemble）\nCV(before threshold_rounder) : 0.389,  CV(after threshold_rounder) : 0.483,  Public LB : 0.463,  Private LB : 0.472\n\n**Changes that didn't work**\n- Optimizing the ensemble weights\nBest result\nCV(before threshold_rounder) : 0.398,  CV(after threshold_rounder) : 0.483,  Public LB : 0.454,  Private LB : 0.463\n- Metric : quadratic weighted kappa\nBest result\nCV(before threshold_rounder) : 0.429,  CV(after threshold_rounder) : 0.485,  Public LB : 0.457,  Private LB : 0.470\n- Using 'PCIAT-PCIAT_Total' as target\nBest result\nCV(before threshold_rounder) : not calculated,  CV(after threshold_rounder) : 0.481,  Public LB : 0.461,  Private LB : 0.471\n\nI hope this post helps you understand the surprising ranking changes.\nThank you for reading !",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3079674": "Hello to all the competition participants. \nThank you to the competition organizers for providing us with this valuable experience.\n\nI gave up trying to catch up with the more skilled participants two months ago.  I am very surprised and confused by the change in ranking after I gave up.  I recognize this result is due to luck, not my ability.\n\nI will share the solution that gave me such unexpected results below.\n\n**Overview**\n- Models\nThree models ensemble (simple average).\nThe models are Vision Transformer (change the input layers code), lightgbm, catboost.\n- Preprocessing\nAggregate parquet data into one row per id using mean, std, etc.\nConvert categorical variables using 'to_dummies (polars)'. \nImpute null values ​​using 'group_by('Basic_Demos-Age', 'Basic_Demos-Sex').mean()' in the training data.\nStandardize features using min_max of the training data.\n- Learning\nMetric : mae\nPerforme cross-validation using 'StratifiedKFold (n_splits=5, y:'sii')'.\nLike many other participants, Use 'threshold_rounder' (after the ensemble).\n- Notebook\n[here](https://www.kaggle.com/code/miyafuru/internet-use-v2?scriptVersionId=201867541)\n\n**Result**\n　CV(before threshold_rounder) : 0.408,  CV(after threshold_rounder) : 0.481,  Public LB : 0.471,  Private LB : 0.476\n\n**Changes that worked**\n- Using the Transformer\nBest result before use（twe models ensemble）\nCV(before threshold_rounder) : 0.389,  CV(after threshold_rounder) : 0.483,  Public LB : 0.463,  Private LB : 0.472\n\n**Changes that didn't work**\n- Optimizing the ensemble weights\nBest result\nCV(before threshold_rounder) : 0.398,  CV(after threshold_rounder) : 0.483,  Public LB : 0.454,  Private LB : 0.463\n- Metric : quadratic weighted kappa\nBest result\nCV(before threshold_rounder) : 0.429,  CV(after threshold_rounder) : 0.485,  Public LB : 0.457,  Private LB : 0.470\n- Using 'PCIAT-PCIAT_Total' as target\nBest result\nCV(before threshold_rounder) : not calculated,  CV(after threshold_rounder) : 0.481,  Public LB : 0.461,  Private LB : 0.471\n\nI hope this post helps you understand the surprising ranking changes.\nThank you for reading !"
  }
}