{
  "id": 552712,
  "title": "2nd Place Writeup",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552712",
  "author_name": "Aradhye Agarwal",
  "post_date": "2024-12-21T06:25:04.796000",
  "votes": 32,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was really surprised to secure 2nd place in this competition, especially since I had essentially quit a few months earlier. My <a href=\"https://www.kaggle.com/code/aradhyeagarwal/starter-notebook-with-polars-gpu\" target=\"_blank\">final approach</a> was based on a fork of the <a href=\"https://www.kaggle.com/code/onodera/starter-notebook-with-polars-gpu\" target=\"_blank\">Starter Notebook</a>. The key observation was that the original notebook didn’t explicitly handle missing values, and it relied on CatBoost to do it automatically. While I wasn’t deeply familiar with CatBoost’s internal approach, I felt that leveraging some domain knowledge through a custom imputation strategy might be a strong alternative.</p>\n<p>Below is the dictionary I used to handle missing values in different columns:</p>\n<pre><code>replacement_strategy = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : \n}\n</code></pre>\n<p>I didn’t leave everything to CatBoost because it can’t <strong>fully “understand”</strong> the meaning behind each feature. For instance, the number of internet usage hours might appear as discrete integers, but we know it has a clear ordering and a specific interpretation (more hours = higher usage). Other features, like IDs of a certain item, also appear as discrete integers but don’t carry such a natural ordering. Only we, as humans, can make that distinction and apply suitable imputation methods.</p>\n<p>Replacement Rules:</p>\n<ul>\n<li>Average: If a feature is numeric (float or integer) and requires an averaged value, the missing entries are replaced by the mean of all existing values.</li>\n<li>New Number: If a feature is integer-based (often treated as categorical) and we want to keep it distinct, we take the current max value in that column and replace missing entries with (max + 1). If the feature is a string-based category, we replace missing entries with \"Null\".</li>\n</ul>\n<p>I only used these strategies if the column was strictly an integer (or recognized as categorical). When a feature was float-based, the “average” replacement was more reasonable, assuming it held a numeric meaning rather than discrete categories.<br>\nI also increased the number of folds for cross-validation from 5 to 20. After that, I saw diminishing returns, so I stopped. Ironically, despite the improved cross-validation score, my public leaderboard score was lower than the baseline notebook, which was quite unexpected. The final private leaderboard result, however, came as a pleasant surprise. I’d love to hear any ideas on why there’s such a significant difference between the public and private scores in the context of my approach.</p>\n<p>Cheers,<br>\nAradhye</p>",
  "messages": [
    {
      "id": 3077579,
      "postDate": "2024-12-21T06:25:04.797Z",
      "content": "<p>I was really surprised to secure 2nd place in this competition, especially since I had essentially quit a few months earlier. My <a href=\"https://www.kaggle.com/code/aradhyeagarwal/starter-notebook-with-polars-gpu\" target=\"_blank\">final approach</a> was based on a fork of the <a href=\"https://www.kaggle.com/code/onodera/starter-notebook-with-polars-gpu\" target=\"_blank\">Starter Notebook</a>. The key observation was that the original notebook didn’t explicitly handle missing values, and it relied on CatBoost to do it automatically. While I wasn’t deeply familiar with CatBoost’s internal approach, I felt that leveraging some domain knowledge through a custom imputation strategy might be a strong alternative.</p>\n<p>Below is the dictionary I used to handle missing values in different columns:</p>\n<pre><code>replacement_strategy = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : \n}\n</code></pre>\n<p>I didn’t leave everything to CatBoost because it can’t <strong>fully “understand”</strong> the meaning behind each feature. For instance, the number of internet usage hours might appear as discrete integers, but we know it has a clear ordering and a specific interpretation (more hours = higher usage). Other features, like IDs of a certain item, also appear as discrete integers but don’t carry such a natural ordering. Only we, as humans, can make that distinction and apply suitable imputation methods.</p>\n<p>Replacement Rules:</p>\n<ul>\n<li>Average: If a feature is numeric (float or integer) and requires an averaged value, the missing entries are replaced by the mean of all existing values.</li>\n<li>New Number: If a feature is integer-based (often treated as categorical) and we want to keep it distinct, we take the current max value in that column and replace missing entries with (max + 1). If the feature is a string-based category, we replace missing entries with \"Null\".</li>\n</ul>\n<p>I only used these strategies if the column was strictly an integer (or recognized as categorical). When a feature was float-based, the “average” replacement was more reasonable, assuming it held a numeric meaning rather than discrete categories.<br>\nI also increased the number of folds for cross-validation from 5 to 20. After that, I saw diminishing returns, so I stopped. Ironically, despite the improved cross-validation score, my public leaderboard score was lower than the baseline notebook, which was quite unexpected. The final private leaderboard result, however, came as a pleasant surprise. I’d love to hear any ideas on why there’s such a significant difference between the public and private scores in the context of my approach.</p>\n<p>Cheers,<br>\nAradhye</p>",
      "rawMarkdown": "I was really surprised to secure 2nd place in this competition, especially since I had essentially quit a few months earlier. My [final approach](https://www.kaggle.com/code/aradhyeagarwal/starter-notebook-with-polars-gpu) was based on a fork of the [Starter Notebook](https://www.kaggle.com/code/onodera/starter-notebook-with-polars-gpu). The key observation was that the original notebook didn’t explicitly handle missing values, and it relied on CatBoost to do it automatically. While I wasn’t deeply familiar with CatBoost’s internal approach, I felt that leveraging some domain knowledge through a custom imputation strategy might be a strong alternative.\n\nBelow is the dictionary I used to handle missing values in different columns:\n\n\n```python\nreplacement_strategy = {\n    'Basic_Demos-Age': 'average',\n    'Basic_Demos-Sex': 'new_number',\n    'CGAS-CGAS_Score': 'average',\n    'Physical-Diastolic_BP': 'average',\n    'Physical-HeartRate': 'average',\n    'Physical-Systolic_BP': 'average',\n    'Fitness_Endurance-Max_Stage': 'average',\n    'Fitness_Endurance-Time_Mins': 'average',\n    'Fitness_Endurance-Time_Sec': 'average',\n    'FGC-FGC_CU': 'new_number',\n    'FGC-FGC_CU_Zone': 'new_number',\n    'FGC-FGC_GSND_Zone': 'new_number',\n    'FGC-FGC_GSD_Zone': 'new_number',\n    'FGC-FGC_PU_Zone': 'new_number',\n    'FGC-FGC_SRL_Zone': 'new_number',\n    'FGC-FGC_SRR_Zone': 'new_number',\n    'FGC-FGC_TL_Zone': 'new_number',\n    'BIA-BIA_Activity_Level_num': 'average',\n    'BIA-BIA_Frame_num': 'average',\n    'SDS-SDS_Total_Raw': 'average',\n    'SDS-SDS_Total_T': 'average',\n    'PreInt_EduHx-computerinternet_hoursday': 'average'\n}\n```\n\nI didn’t leave everything to CatBoost because it can’t **fully “understand”** the meaning behind each feature. For instance, the number of internet usage hours might appear as discrete integers, but we know it has a clear ordering and a specific interpretation (more hours = higher usage). Other features, like IDs of a certain item, also appear as discrete integers but don’t carry such a natural ordering. Only we, as humans, can make that distinction and apply suitable imputation methods.\n\n Replacement Rules:\n\n- Average: If a feature is numeric (float or integer) and requires an averaged value, the missing entries are replaced by the mean of all existing values.\n- New Number: If a feature is integer-based (often treated as categorical) and we want to keep it distinct, we take the current max value in that column and replace missing entries with (max + 1). If the feature is a string-based category, we replace missing entries with \"Null\".\n\nI only used these strategies if the column was strictly an integer (or recognized as categorical). When a feature was float-based, the “average” replacement was more reasonable, assuming it held a numeric meaning rather than discrete categories.\nI also increased the number of folds for cross-validation from 5 to 20. After that, I saw diminishing returns, so I stopped. Ironically, despite the improved cross-validation score, my public leaderboard score was lower than the baseline notebook, which was quite unexpected. The final private leaderboard result, however, came as a pleasant surprise. I’d love to hear any ideas on why there’s such a significant difference between the public and private scores in the context of my approach.\n\nCheers,\nAradhye\n\n",
      "votes": 32
    },
    {
      "id": 3242602,
      "postDate": "2025-07-06T06:05:31.393Z",
      "content": "<p>Excellent work, great write up too.</p>",
      "rawMarkdown": "Excellent work, great write up too.",
      "votes": 3
    },
    {
      "id": 3242446,
      "postDate": "2025-07-05T21:41:30.820Z",
      "content": "<p>hi, great work! i am a newbie in ml so i got a simple question: i've seen that for missing values imputation you could run a random forest and work with the istances'distance since the model \"fully\" uderstands the data's distribution , what would you consider more important, your feature's domain understanding or whatever the random forest technique proposes?</p>",
      "rawMarkdown": "hi, great work! i am a newbie in ml so i got a simple question: i've seen that for missing values imputation you could run a random forest and work with the istances'distance since the model \"fully\" uderstands the data's distribution , what would you consider more important, your feature's domain understanding or whatever the random forest technique proposes?",
      "votes": 3
    },
    {
      "id": 3078440,
      "postDate": "2024-12-22T10:10:45.410Z",
      "content": "<p>Congrats, Aradhye! Impressive how your thoughtful approach made a big impact—well deserved!</p>",
      "rawMarkdown": "Congrats, Aradhye! Impressive how your thoughtful approach made a big impact—well deserved!",
      "votes": 3,
      "replies": [
        {
          "id": 3078476,
          "postDate": "2024-12-22T11:19:27.187Z",
          "content": "<p>Thank you Hamed!</p>",
          "rawMarkdown": "Thank you Hamed!",
          "votes": 2
        }
      ]
    },
    {
      "id": 3077594,
      "postDate": "2024-12-21T06:41:42.877Z",
      "content": "<p>It was a very surprising competition. Thank you for sharing, congratulations.</p>",
      "rawMarkdown": "It was a very surprising competition. Thank you for sharing, congratulations.",
      "votes": 3
    },
    {
      "id": 3079615,
      "postDate": "2024-12-23T22:42:55.530Z",
      "content": "<p>Wow! Great work!</p>",
      "rawMarkdown": "Wow! Great work!",
      "votes": 1
    },
    {
      "id": 3079169,
      "postDate": "2024-12-23T09:35:04.990Z",
      "content": "<p>Congrats! Thank you for sharing approach!</p>",
      "rawMarkdown": "Congrats! Thank you for sharing approach!",
      "votes": 2
    },
    {
      "id": 3077630,
      "postDate": "2024-12-21T08:13:37.240Z",
      "content": "<p>What's public leaderboard score for this notebook?</p>",
      "rawMarkdown": "What's public leaderboard score for this notebook?",
      "votes": 2,
      "replies": [
        {
          "id": 3077658,
          "postDate": "2024-12-21T08:51:46.590Z",
          "content": "<p>It was 0.435</p>",
          "rawMarkdown": "It was 0.435",
          "votes": 2
        }
      ]
    },
    {
      "id": 3081180,
      "postDate": "2024-12-26T11:28:08.790Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3242602,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-07-06T06:05:31.393000",
      "content": "<p>Excellent work, great write up too.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3242446,
      "author_name": "Giacomo Zuccolotto",
      "author_url": "",
      "post_date": "2025-07-05T21:41:30.820000",
      "content": "<p>hi, great work! i am a newbie in ml so i got a simple question: i've seen that for missing values imputation you could run a random forest and work with the istances'distance since the model \"fully\" uderstands the data's distribution , what would you consider more important, your feature's domain understanding or whatever the random forest technique proposes?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3078440,
      "author_name": "Hamed Abedi",
      "author_url": "",
      "post_date": "2024-12-22T10:10:45.410000",
      "content": "<p>Congrats, Aradhye! Impressive how your thoughtful approach made a big impact—well deserved!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3078476,
          "author_name": "Aradhye Agarwal",
          "author_url": "",
          "post_date": "2024-12-22T11:19:27.187000",
          "content": "<p>Thank you Hamed!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3077594,
      "author_name": "Sems Kurtoglu",
      "author_url": "",
      "post_date": "2024-12-21T06:41:42.877000",
      "content": "<p>It was a very surprising competition. Thank you for sharing, congratulations.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3079615,
      "author_name": "BobbyData123",
      "author_url": "",
      "post_date": "2024-12-23T22:42:55.530000",
      "content": "<p>Wow! Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3079169,
      "author_name": "Ivan Lebed",
      "author_url": "",
      "post_date": "2024-12-23T09:35:04.990000",
      "content": "<p>Congrats! Thank you for sharing approach!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3077630,
      "author_name": "Reki",
      "author_url": "",
      "post_date": "2024-12-21T08:13:37.240000",
      "content": "<p>What's public leaderboard score for this notebook?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3077658,
          "author_name": "Aradhye Agarwal",
          "author_url": "",
          "post_date": "2024-12-21T08:51:46.590000",
          "content": "<p>It was 0.435</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3081180,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-26T11:28:08.790000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3077579": "I was really surprised to secure 2nd place in this competition, especially since I had essentially quit a few months earlier. My [final approach](https://www.kaggle.com/code/aradhyeagarwal/starter-notebook-with-polars-gpu) was based on a fork of the [Starter Notebook](https://www.kaggle.com/code/onodera/starter-notebook-with-polars-gpu). The key observation was that the original notebook didn’t explicitly handle missing values, and it relied on CatBoost to do it automatically. While I wasn’t deeply familiar with CatBoost’s internal approach, I felt that leveraging some domain knowledge through a custom imputation strategy might be a strong alternative.\n\nBelow is the dictionary I used to handle missing values in different columns:\n\n\n```python\nreplacement_strategy = {\n    'Basic_Demos-Age': 'average',\n    'Basic_Demos-Sex': 'new_number',\n    'CGAS-CGAS_Score': 'average',\n    'Physical-Diastolic_BP': 'average',\n    'Physical-HeartRate': 'average',\n    'Physical-Systolic_BP': 'average',\n    'Fitness_Endurance-Max_Stage': 'average',\n    'Fitness_Endurance-Time_Mins': 'average',\n    'Fitness_Endurance-Time_Sec': 'average',\n    'FGC-FGC_CU': 'new_number',\n    'FGC-FGC_CU_Zone': 'new_number',\n    'FGC-FGC_GSND_Zone': 'new_number',\n    'FGC-FGC_GSD_Zone': 'new_number',\n    'FGC-FGC_PU_Zone': 'new_number',\n    'FGC-FGC_SRL_Zone': 'new_number',\n    'FGC-FGC_SRR_Zone': 'new_number',\n    'FGC-FGC_TL_Zone': 'new_number',\n    'BIA-BIA_Activity_Level_num': 'average',\n    'BIA-BIA_Frame_num': 'average',\n    'SDS-SDS_Total_Raw': 'average',\n    'SDS-SDS_Total_T': 'average',\n    'PreInt_EduHx-computerinternet_hoursday': 'average'\n}\n```\n\nI didn’t leave everything to CatBoost because it can’t **fully “understand”** the meaning behind each feature. For instance, the number of internet usage hours might appear as discrete integers, but we know it has a clear ordering and a specific interpretation (more hours = higher usage). Other features, like IDs of a certain item, also appear as discrete integers but don’t carry such a natural ordering. Only we, as humans, can make that distinction and apply suitable imputation methods.\n\n Replacement Rules:\n\n- Average: If a feature is numeric (float or integer) and requires an averaged value, the missing entries are replaced by the mean of all existing values.\n- New Number: If a feature is integer-based (often treated as categorical) and we want to keep it distinct, we take the current max value in that column and replace missing entries with (max + 1). If the feature is a string-based category, we replace missing entries with \"Null\".\n\nI only used these strategies if the column was strictly an integer (or recognized as categorical). When a feature was float-based, the “average” replacement was more reasonable, assuming it held a numeric meaning rather than discrete categories.\nI also increased the number of folds for cross-validation from 5 to 20. After that, I saw diminishing returns, so I stopped. Ironically, despite the improved cross-validation score, my public leaderboard score was lower than the baseline notebook, which was quite unexpected. The final private leaderboard result, however, came as a pleasant surprise. I’d love to hear any ideas on why there’s such a significant difference between the public and private scores in the context of my approach.\n\nCheers,\nAradhye\n\n",
    "3242602": "Excellent work, great write up too.",
    "3242446": "hi, great work! i am a newbie in ml so i got a simple question: i've seen that for missing values imputation you could run a random forest and work with the istances'distance since the model \"fully\" uderstands the data's distribution , what would you consider more important, your feature's domain understanding or whatever the random forest technique proposes?",
    "3078440": "Congrats, Aradhye! Impressive how your thoughtful approach made a big impact—well deserved!",
    "3077594": "It was a very surprising competition. Thank you for sharing, congratulations.",
    "3079615": "Wow! Great work!",
    "3079169": "Congrats! Thank you for sharing approach!",
    "3077630": "What's public leaderboard score for this notebook?",
    "3081180": ""
  }
}