{
  "id": 552534,
  "title": "Classification kinda approach that didn't work",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552534",
  "author_name": "",
  "post_date": "2024-12-20T07:09:19.308405100Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>My approach was little bit different than public notebooks so I wanted to share it. I cleaned the target like this which made sense to me.</p>\n<ul>\n<li>Drop rows that didn't take PCIAT</li>\n<li>Drop one row that didn't answer any of the questions but took the test</li>\n<li>Fill rest of the rows with row median</li>\n</ul>\n<pre><code>pciat_question_columns = [column  column  df.columns.tolist()  column.startswith()][:-]\npciat_total_column = \n\n\ndf = df.loc[df[pciat_total_column].notna()].reset_index(drop=)\n\n\ndf[] = df[pciat_question_columns].isnull().(axis=)\ndf = df.loc[df[] != ].reset_index(drop=)\n\n\ndf[] = df[pciat_question_columns].median(axis=)\n idx  (df.shape[]):\n    df.loc[idx, pciat_question_columns] = df.loc[idx, pciat_question_columns].fillna(df.loc[idx, ])\n</code></pre>\n<p>Then I recalculated PCIAT total and squashed continuous values between 0 and 1, and trained using cross entropy loss.</p>\n<pre><code>df[] = df[pciat_question_columns].(axis=)\ndf[] = df[] / \n</code></pre>\n<p>My best single model oof score was 0.4882 and public lb score was 0.45x, but it scored 0.39x on private lb. Congrats to all winners!</p>",
  "messages": [
    {
      "id": "3076669",
      "postDate": "12/20/2024 07:09:19",
      "content": "<p>My approach was little bit different than public notebooks so I wanted to share it. I cleaned the target like this which made sense to me.</p>\n<ul>\n<li>Drop rows that didn't take PCIAT</li>\n<li>Drop one row that didn't answer any of the questions but took the test</li>\n<li>Fill rest of the rows with row median</li>\n</ul>\n<pre><code>pciat_question_columns = [column  column  df.columns.tolist()  column.startswith()][:-]\npciat_total_column = \n\n\ndf = df.loc[df[pciat_total_column].notna()].reset_index(drop=)\n\n\ndf[] = df[pciat_question_columns].isnull().(axis=)\ndf = df.loc[df[] != ].reset_index(drop=)\n\n\ndf[] = df[pciat_question_columns].median(axis=)\n idx  (df.shape[]):\n    df.loc[idx, pciat_question_columns] = df.loc[idx, pciat_question_columns].fillna(df.loc[idx, ])\n</code></pre>\n<p>Then I recalculated PCIAT total and squashed continuous values between 0 and 1, and trained using cross entropy loss.</p>\n<pre><code>df[] = df[pciat_question_columns].(axis=)\ndf[] = df[] / \n</code></pre>\n<p>My best single model oof score was 0.4882 and public lb score was 0.45x, but it scored 0.39x on private lb. Congrats to all winners!</p>",
      "rawMarkdown": "My approach was little bit different than public notebooks so I wanted to share it. I cleaned the target like this which made sense to me.\n\n* Drop rows that didn't take PCIAT\n* Drop one row that didn't answer any of the questions but took the test\n* Fill rest of the rows with row median\n```python\npciat_question_columns = [column for column in df.columns.tolist() if column.startswith('PCIAT')][1:-1]\npciat_total_column = 'PCIAT-PCIAT_Total'\n\n# Drop rows that didn't take the test\ndf = df.loc[df[pciat_total_column].notna()].reset_index(drop=True)\n\n# Drop one row that didn't answer any of the questions\ndf['pciat_missing_count'] = df[pciat_question_columns].isnull().sum(axis=1)\ndf = df.loc[df['pciat_missing_count'] != 20].reset_index(drop=True)\n\n# Fill missing values answers with median of the row\ndf['pciat_median'] = df[pciat_question_columns].median(axis=1)\nfor idx in range(df.shape[0]):\n    df.loc[idx, pciat_question_columns] = df.loc[idx, pciat_question_columns].fillna(df.loc[idx, 'pciat_median'])\n```\nThen I recalculated PCIAT total and squashed continuous values between 0 and 1, and trained using cross entropy loss.\n```python\ndf['pciat_total'] = df[pciat_question_columns].sum(axis=1)\ndf['pciat_total_normalized'] = df['pciat_total'] / 100.\n```\n\nMy best single model oof score was 0.4882 and public lb score was 0.45x, but it scored 0.39x on private lb. Congrats to all winners!",
      "votes": null
    },
    {
      "id": "3076679",
      "postDate": "12/20/2024 07:19:19",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> thanks for sharing the approach. Regression models worked here. I to tried to use a classifier early on and it did not work, soi I switched to regression and did not change my mind thereafter.</p>",
      "rawMarkdown": "gunesevitan thanks for sharing the approach. Regression models worked here. I to tried to use a classifier early on and it did not work, soi I switched to regression and did not change my mind thereafter.",
      "votes": null
    },
    {
      "id": "3076692",
      "postDate": "12/20/2024 07:33:32",
      "content": "<p>Thanks for sharing! good work!</p>",
      "rawMarkdown": "Thanks for sharing! good work!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076679,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/20/2024 07:19:19",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> thanks for sharing the approach. Regression models worked here. I to tried to use a classifier early on and it did not work, soi I switched to regression and did not change my mind thereafter.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076692,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/20/2024 07:33:32",
      "content": "<p>Thanks for sharing! good work!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3076669": "My approach was little bit different than public notebooks so I wanted to share it. I cleaned the target like this which made sense to me.\n\n* Drop rows that didn't take PCIAT\n* Drop one row that didn't answer any of the questions but took the test\n* Fill rest of the rows with row median\n```python\npciat_question_columns = [column for column in df.columns.tolist() if column.startswith('PCIAT')][1:-1]\npciat_total_column = 'PCIAT-PCIAT_Total'\n\n# Drop rows that didn't take the test\ndf = df.loc[df[pciat_total_column].notna()].reset_index(drop=True)\n\n# Drop one row that didn't answer any of the questions\ndf['pciat_missing_count'] = df[pciat_question_columns].isnull().sum(axis=1)\ndf = df.loc[df['pciat_missing_count'] != 20].reset_index(drop=True)\n\n# Fill missing values answers with median of the row\ndf['pciat_median'] = df[pciat_question_columns].median(axis=1)\nfor idx in range(df.shape[0]):\n    df.loc[idx, pciat_question_columns] = df.loc[idx, pciat_question_columns].fillna(df.loc[idx, 'pciat_median'])\n```\nThen I recalculated PCIAT total and squashed continuous values between 0 and 1, and trained using cross entropy loss.\n```python\ndf['pciat_total'] = df[pciat_question_columns].sum(axis=1)\ndf['pciat_total_normalized'] = df['pciat_total'] / 100.\n```\n\nMy best single model oof score was 0.4882 and public lb score was 0.45x, but it scored 0.39x on private lb. Congrats to all winners!",
    "3076679": "gunesevitan thanks for sharing the approach. Regression models worked here. I to tried to use a classifier early on and it did not work, soi I switched to regression and did not change my mind thereafter.",
    "3076692": "Thanks for sharing! good work!"
  },
  "source": "meta"
}