{
  "id": 535121,
  "title": "Starter Notebook: Multi-Target Prediction Using CatBoost",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535121",
  "author_name": "",
  "post_date": "2024-09-20T11:41:54.602861500Z",
  "votes": 36,
  "comment_count": 8,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/tubotubo/starter-notebook-multi-target-prediction?scriptVersionId=197454913\" target=\"_blank\">Starter Notebook: Multi-Target Prediction Using CatBoost</a></p>\n<h2>Introduction</h2>\n<p>In this notebook, we demonstrate how to use CatBoost's MultiRegression to predict multiple targets. The goal is to build a baseline model that can be useful for Kaggle competitions.</p>\n<h2>What We Did</h2>\n<ol>\n<li><p><strong>Multi-Target Prediction Using CatBoost's MultiRegression</strong>  </p>\n<ul>\n<li><strong>Features</strong>: We only used the features from the CSV file and did not use any from the parquet file.  </li>\n<li><strong>Data Preprocessing</strong>: Rows with missing target values were excluded from the training data.</li></ul></li>\n<li><p><strong>Converting PCIAT-PCIAT_Total Predictions to Sii Predictions</strong>  </p>\n<ul>\n<li>Based on the predicted values, we converted the predictions to Sii using a threshold.  </li>\n<li>The threshold was optimized using <a href=\"https://optuna.org/\" target=\"_blank\">Optuna</a>.</li></ul></li>\n</ol>\n<h2>Result</h2>\n<ul>\n<li>CV: 0.461</li>\n<li>LB: 0.441</li>\n</ul>\n<h2>Future Directions</h2>\n<ul>\n<li><p><strong>Utilizing Features from the Parquet File</strong>  <br>\nWe plan to incorporate features from the parquet files to further improve the model's performance.</p></li>\n<li><p><strong>Pseudo-Labeling</strong>  <br>\nWe aim to implement pseudo-labeling to improve performance when labeled data is scarce.</p></li>\n<li><p><strong>Feature Engineering</strong>  <br>\nWe will explore feature combinations and transformations to enhance the model's accuracy.</p></li>\n</ul>\n<p>This notebook serves as a useful baseline for multi-target prediction, especially for Kaggle participants. We plan to update it with improvements and new ideas in the future. Please feel free to share your feedback or suggestions!</p>",
  "messages": [
    {
      "id": "2993939",
      "postDate": "09/20/2024 11:41:54",
      "content": "<p><a href=\"https://www.kaggle.com/code/tubotubo/starter-notebook-multi-target-prediction?scriptVersionId=197454913\" target=\"_blank\">Starter Notebook: Multi-Target Prediction Using CatBoost</a></p>\n<h2>Introduction</h2>\n<p>In this notebook, we demonstrate how to use CatBoost's MultiRegression to predict multiple targets. The goal is to build a baseline model that can be useful for Kaggle competitions.</p>\n<h2>What We Did</h2>\n<ol>\n<li><p><strong>Multi-Target Prediction Using CatBoost's MultiRegression</strong>  </p>\n<ul>\n<li><strong>Features</strong>: We only used the features from the CSV file and did not use any from the parquet file.  </li>\n<li><strong>Data Preprocessing</strong>: Rows with missing target values were excluded from the training data.</li></ul></li>\n<li><p><strong>Converting PCIAT-PCIAT_Total Predictions to Sii Predictions</strong>  </p>\n<ul>\n<li>Based on the predicted values, we converted the predictions to Sii using a threshold.  </li>\n<li>The threshold was optimized using <a href=\"https://optuna.org/\" target=\"_blank\">Optuna</a>.</li></ul></li>\n</ol>\n<h2>Result</h2>\n<ul>\n<li>CV: 0.461</li>\n<li>LB: 0.441</li>\n</ul>\n<h2>Future Directions</h2>\n<ul>\n<li><p><strong>Utilizing Features from the Parquet File</strong>  <br>\nWe plan to incorporate features from the parquet files to further improve the model's performance.</p></li>\n<li><p><strong>Pseudo-Labeling</strong>  <br>\nWe aim to implement pseudo-labeling to improve performance when labeled data is scarce.</p></li>\n<li><p><strong>Feature Engineering</strong>  <br>\nWe will explore feature combinations and transformations to enhance the model's accuracy.</p></li>\n</ul>\n<p>This notebook serves as a useful baseline for multi-target prediction, especially for Kaggle participants. We plan to update it with improvements and new ideas in the future. Please feel free to share your feedback or suggestions!</p>",
      "rawMarkdown": "[Starter Notebook: Multi-Target Prediction Using CatBoost](https://www.kaggle.com/code/tubotubo/starter-notebook-multi-target-prediction?scriptVersionId=197454913)\n\n## Introduction\nIn this notebook, we demonstrate how to use CatBoost's MultiRegression to predict multiple targets. The goal is to build a baseline model that can be useful for Kaggle competitions.\n\n## What We Did\n\n1. **Multi-Target Prediction Using CatBoost's MultiRegression**  \n   - **Features**: We only used the features from the CSV file and did not use any from the parquet file.  \n   - **Data Preprocessing**: Rows with missing target values were excluded from the training data.\n\n2. **Converting PCIAT-PCIAT_Total Predictions to Sii Predictions**  \n   - Based on the predicted values, we converted the predictions to Sii using a threshold.  \n   - The threshold was optimized using [Optuna](https://optuna.org/).\n\n## Result\n\n- CV: 0.461\n- LB: 0.441\n\n## Future Directions\n\n- **Utilizing Features from the Parquet File**  \n  We plan to incorporate features from the parquet files to further improve the model's performance.\n\n- **Pseudo-Labeling**  \n  We aim to implement pseudo-labeling to improve performance when labeled data is scarce.\n\n- **Feature Engineering**  \n  We will explore feature combinations and transformations to enhance the model's accuracy.\n\nThis notebook serves as a useful baseline for multi-target prediction, especially for Kaggle participants. We plan to update it with improvements and new ideas in the future. Please feel free to share your feedback or suggestions!",
      "votes": null
    },
    {
      "id": "2994301",
      "postDate": "09/20/2024 18:43:46",
      "content": "<p>Problem is MultiClass  Classification ? or MultiClass Regression. </p>\n<p>my CV : 0.3299 | LB : 0.345</p>",
      "rawMarkdown": "Problem is MultiClass  Classification ? or MultiClass Regression. \n\nmy CV : 0.3299 | LB : 0.345",
      "votes": null
    },
    {
      "id": "2994416",
      "postDate": "09/20/2024 22:38:04",
      "content": "<p>This is a tricky problem! You can try either approach <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
      "rawMarkdown": "This is a tricky problem! You can try either approach @abdmental01",
      "votes": null
    },
    {
      "id": "2994422",
      "postDate": "09/20/2024 22:54:42",
      "content": "<p>Can you Suggest Which one is more Suitable?</p>",
      "rawMarkdown": "Can you Suggest Which one is more Suitable?",
      "votes": null
    },
    {
      "id": "2996156",
      "postDate": "09/23/2024 06:20:27",
      "content": "<p>It would be great if you can suggest upon how to utilize features from the Parquet file. It is very difficult for me.</p>",
      "rawMarkdown": "It would be great if you can suggest upon how to utilize features from the Parquet file. It is very difficult for me.",
      "votes": null
    },
    {
      "id": "2996510",
      "postDate": "09/23/2024 15:08:52",
      "content": "<p>Given there's only four classes, I'd stick with classification from my end at first. However, I think people have had great success reconfiguring these problems as a regression task. Specifically, I'm thinking of the Common Lit prizes.</p>\n<p>I think this has to do with the fact the labels in this case are ordered, meaning a score of 1 is greater than a score of 0 and score of 2 greater than 1, etc. This means the labels have an inherent order to them and might help framing it as regression.</p>\n<p>But as the case always is, do both, experiment, and pick what works best.</p>",
      "rawMarkdown": "Given there's only four classes, I'd stick with classification from my end at first. However, I think people have had great success reconfiguring these problems as a regression task. Specifically, I'm thinking of the Common Lit prizes.\n\nI think this has to do with the fact the labels in this case are ordered, meaning a score of 1 is greater than a score of 0 and score of 2 greater than 1, etc. This means the labels have an inherent order to them and might help framing it as regression.\n\nBut as the case always is, do both, experiment, and pick what works best.",
      "votes": null
    },
    {
      "id": "2996884",
      "postDate": "09/24/2024 03:12:22",
      "content": "<p>Thanks, Mate .</p>",
      "rawMarkdown": "Thanks, Mate .",
      "votes": null
    },
    {
      "id": "3015681",
      "postDate": "10/12/2024 18:48:32",
      "content": "<p>is true that lightgbm, catboost, xgboost has upper uper hand over deep learning such pytorch in tabular data such as in this competition (most public notebboks use lightgbm , catboost in their approch)?</p>",
      "rawMarkdown": "is true that lightgbm, catboost, xgboost has upper uper hand over deep learning such pytorch in tabular data such as in this competition (most public notebboks use lightgbm , catboost in their approch)?",
      "votes": null
    },
    {
      "id": "3029874",
      "postDate": "10/27/2024 21:09:32",
      "content": "<p>This is a great starting point.</p>",
      "rawMarkdown": "This is a great starting point.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2994301,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "09/20/2024 18:43:46",
      "content": "<p>Problem is MultiClass  Classification ? or MultiClass Regression. </p>\n<p>my CV : 0.3299 | LB : 0.345</p>",
      "votes": null,
      "replies": [
        {
          "id": 2994416,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "09/20/2024 22:38:04",
          "content": "<p>This is a tricky problem! You can try either approach <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 2994422,
              "author_name": "abdmental01",
              "author_url": "",
              "post_date": "09/20/2024 22:54:42",
              "content": "<p>Can you Suggest Which one is more Suitable?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2996510,
                  "author_name": "msthil",
                  "author_url": "",
                  "post_date": "09/23/2024 15:08:52",
                  "content": "<p>Given there's only four classes, I'd stick with classification from my end at first. However, I think people have had great success reconfiguring these problems as a regression task. Specifically, I'm thinking of the Common Lit prizes.</p>\n<p>I think this has to do with the fact the labels in this case are ordered, meaning a score of 1 is greater than a score of 0 and score of 2 greater than 1, etc. This means the labels have an inherent order to them and might help framing it as regression.</p>\n<p>But as the case always is, do both, experiment, and pick what works best.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2996156,
      "author_name": "akihiroorita",
      "author_url": "",
      "post_date": "09/23/2024 06:20:27",
      "content": "<p>It would be great if you can suggest upon how to utilize features from the Parquet file. It is very difficult for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2996884,
      "author_name": "k2001g",
      "author_url": "",
      "post_date": "09/24/2024 03:12:22",
      "content": "<p>Thanks, Mate .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3015681,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "10/12/2024 18:48:32",
      "content": "<p>is true that lightgbm, catboost, xgboost has upper uper hand over deep learning such pytorch in tabular data such as in this competition (most public notebboks use lightgbm , catboost in their approch)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3029874,
      "author_name": "wekeed",
      "author_url": "",
      "post_date": "10/27/2024 21:09:32",
      "content": "<p>This is a great starting point.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2993939": "[Starter Notebook: Multi-Target Prediction Using CatBoost](https://www.kaggle.com/code/tubotubo/starter-notebook-multi-target-prediction?scriptVersionId=197454913)\n\n## Introduction\nIn this notebook, we demonstrate how to use CatBoost's MultiRegression to predict multiple targets. The goal is to build a baseline model that can be useful for Kaggle competitions.\n\n## What We Did\n\n1. **Multi-Target Prediction Using CatBoost's MultiRegression**  \n   - **Features**: We only used the features from the CSV file and did not use any from the parquet file.  \n   - **Data Preprocessing**: Rows with missing target values were excluded from the training data.\n\n2. **Converting PCIAT-PCIAT_Total Predictions to Sii Predictions**  \n   - Based on the predicted values, we converted the predictions to Sii using a threshold.  \n   - The threshold was optimized using [Optuna](https://optuna.org/).\n\n## Result\n\n- CV: 0.461\n- LB: 0.441\n\n## Future Directions\n\n- **Utilizing Features from the Parquet File**  \n  We plan to incorporate features from the parquet files to further improve the model's performance.\n\n- **Pseudo-Labeling**  \n  We aim to implement pseudo-labeling to improve performance when labeled data is scarce.\n\n- **Feature Engineering**  \n  We will explore feature combinations and transformations to enhance the model's accuracy.\n\nThis notebook serves as a useful baseline for multi-target prediction, especially for Kaggle participants. We plan to update it with improvements and new ideas in the future. Please feel free to share your feedback or suggestions!",
    "2994301": "Problem is MultiClass  Classification ? or MultiClass Regression. \n\nmy CV : 0.3299 | LB : 0.345",
    "2994416": "This is a tricky problem! You can try either approach @abdmental01",
    "2994422": "Can you Suggest Which one is more Suitable?",
    "2996156": "It would be great if you can suggest upon how to utilize features from the Parquet file. It is very difficult for me.",
    "2996510": "Given there's only four classes, I'd stick with classification from my end at first. However, I think people have had great success reconfiguring these problems as a regression task. Specifically, I'm thinking of the Common Lit prizes.\n\nI think this has to do with the fact the labels in this case are ordered, meaning a score of 1 is greater than a score of 0 and score of 2 greater than 1, etc. This means the labels have an inherent order to them and might help framing it as regression.\n\nBut as the case always is, do both, experiment, and pick what works best.",
    "2996884": "Thanks, Mate .",
    "3015681": "is true that lightgbm, catboost, xgboost has upper uper hand over deep learning such pytorch in tabular data such as in this competition (most public notebboks use lightgbm , catboost in their approch)?",
    "3029874": "This is a great starting point."
  },
  "source": "meta"
}