{
  "id": 340932,
  "title": "Trying to get to .6",
  "url": "/competitions/amex-default-prediction/discussion/340932",
  "author_name": "Campbell Hutcheson",
  "post_date": "2022-07-31T14:38:39.296000",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>So, I'm trying to get to .6 without using any of the publicly available code as a practice. I'm at .451 and wondered if anyone might have any advice.</p>\n<p>1) After creating CV folds, I cut down the dataset to just the following columns to make my models faster to train and the data easier to explore.</p>\n<p>I picked the following columns just based upon single column predictiveness of the target.</p>\n<p>[\"customer_ID\", \"P_2\", \"B_18\", \"B_9\", \"B_2\", \"D_48\", \"D_44\", \"B_3\", \"B_1\", \"B_37\", \"D_75\", \"B_11\", \"B_19\", \"B_22\", \"D_61\", \"B_6\", \"B_7\", \"B_23\", \"D_62\", \"B_38\", \"D_55\"]</p>\n<p>My intuition is that the number of columns that actually matter must be very small and most of the columns must be mostly irrelevant.</p>\n<p>2) I threw the data into Catboost as a first attempt and used the scale_pos_weight with a 4:1 weighting to reflect the imbalance of the data. </p>\n<p>I limited the model to a small number of iterations and small tree size because I expect that there is a real risk of overfitting.</p>\n<p>3) I did predictions on the test data rows and then did a majority vote per customer ID to reach a prediction.</p>\n<p>What do you think would be the best next steps?</p>\n<p>1) Is it possible to hit .6 with just the features that I picked or should my next step be to add more features?</p>\n<p>2) I could just a hyper parameter optimization tool like Optuna or Skopt but I feel that with my results as low as they are, there is probably something fundamentally wrong that I need to fix first.</p>",
  "messages": [
    {
      "id": 1878670,
      "postDate": "2022-07-31T14:38:39.297Z",
      "content": "<p>So, I'm trying to get to .6 without using any of the publicly available code as a practice. I'm at .451 and wondered if anyone might have any advice.</p>\n<p>1) After creating CV folds, I cut down the dataset to just the following columns to make my models faster to train and the data easier to explore.</p>\n<p>I picked the following columns just based upon single column predictiveness of the target.</p>\n<p>[\"customer_ID\", \"P_2\", \"B_18\", \"B_9\", \"B_2\", \"D_48\", \"D_44\", \"B_3\", \"B_1\", \"B_37\", \"D_75\", \"B_11\", \"B_19\", \"B_22\", \"D_61\", \"B_6\", \"B_7\", \"B_23\", \"D_62\", \"B_38\", \"D_55\"]</p>\n<p>My intuition is that the number of columns that actually matter must be very small and most of the columns must be mostly irrelevant.</p>\n<p>2) I threw the data into Catboost as a first attempt and used the scale_pos_weight with a 4:1 weighting to reflect the imbalance of the data. </p>\n<p>I limited the model to a small number of iterations and small tree size because I expect that there is a real risk of overfitting.</p>\n<p>3) I did predictions on the test data rows and then did a majority vote per customer ID to reach a prediction.</p>\n<p>What do you think would be the best next steps?</p>\n<p>1) Is it possible to hit .6 with just the features that I picked or should my next step be to add more features?</p>\n<p>2) I could just a hyper parameter optimization tool like Optuna or Skopt but I feel that with my results as low as they are, there is probably something fundamentally wrong that I need to fix first.</p>",
      "rawMarkdown": "So, I'm trying to get to .6 without using any of the publicly available code as a practice. I'm at .451 and wondered if anyone might have any advice.\n\n1) After creating CV folds, I cut down the dataset to just the following columns to make my models faster to train and the data easier to explore.\n\nI picked the following columns just based upon single column predictiveness of the target.\n\n[\"customer_ID\", \"P_2\", \"B_18\", \"B_9\", \"B_2\", \"D_48\", \"D_44\", \"B_3\", \"B_1\", \"B_37\", \"D_75\", \"B_11\", \"B_19\", \"B_22\", \"D_61\", \"B_6\", \"B_7\", \"B_23\", \"D_62\", \"B_38\", \"D_55\"]\n\nMy intuition is that the number of columns that actually matter must be very small and most of the columns must be mostly irrelevant.\n\n2) I threw the data into Catboost as a first attempt and used the scale_pos_weight with a 4:1 weighting to reflect the imbalance of the data. \n\nI limited the model to a small number of iterations and small tree size because I expect that there is a real risk of overfitting.\n\n3) I did predictions on the test data rows and then did a majority vote per customer ID to reach a prediction.\n\nWhat do you think would be the best next steps?\n\n1) Is it possible to hit .6 with just the features that I picked or should my next step be to add more features?\n\n2) I could just a hyper parameter optimization tool like Optuna or Skopt but I feel that with my results as low as they are, there is probably something fundamentally wrong that I need to fix first.",
      "votes": 11
    },
    {
      "id": 1878860,
      "postDate": "2022-07-31T17:16:05.440Z",
      "content": "<p>I'd suggest these next steps:</p>\n<ol>\n<li>As <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> recommends: Keep only the last row for every customer (i.e. the most recent data) and get rid of the majority vote.</li>\n<li>Don't submit the binary labels, but the predicted probabilities (CatBoostClassifier.predict_proba). The competition's evaluation metric needs the probabilities. (You'd submit the labels if the competition were evaluated by accuracy.)</li>\n<li>Don't be afraid of overfitting: Use large trees and unlimited iterations with early stopping.</li>\n<li>Add some more features.</li>\n<li>Play with the catboost parameter colsample_bylevel. You don't need Optuna for that.</li>\n</ol>\n<p>And tell us about the results…</p>",
      "rawMarkdown": "I'd suggest these next steps:\n1. As @lucasmorin recommends: Keep only the last row for every customer (i.e. the most recent data) and get rid of the majority vote.\n2. Don't submit the binary labels, but the predicted probabilities (CatBoostClassifier.predict_proba). The competition's evaluation metric needs the probabilities. (You'd submit the labels if the competition were evaluated by accuracy.)\n3. Don't be afraid of overfitting: Use large trees and unlimited iterations with early stopping.\n4. Add some more features.\n5. Play with the catboost parameter colsample_bylevel. You don't need Optuna for that.\n\nAnd tell us about the results...",
      "votes": 6,
      "replies": [
        {
          "id": 1880105,
          "postDate": "2022-08-01T13:17:23.977Z",
          "content": "<p>I still struggle with 3). Would you recommend this outside of kaggle ? </p>",
          "rawMarkdown": "I still struggle with 3). Would you recommend this outside of kaggle ? "
        },
        {
          "id": 1880125,
          "postDate": "2022-08-01T13:30:45.217Z",
          "content": "<p>Great advice! </p>\n<p>I was able to get to .741 : O using just 1) and 2), so now I'll take a look at 3), 4) and 5)!</p>\n<p>I am using InterpretML to try to get an idea of feature interactions, I don't see anyone else using it on here, though.</p>\n<p>Do folks here normally look at t-SNE plots to try to get an idea of feature clusters (to figure out which features to add)? Or maybe do hierarchical clustering?</p>",
          "rawMarkdown": "Great advice! \n\nI was able to get to .741 : O using just 1) and 2), so now I'll take a look at 3), 4) and 5)!\n\nI am using InterpretML to try to get an idea of feature interactions, I don't see anyone else using it on here, though.\n\nDo folks here normally look at t-SNE plots to try to get an idea of feature clusters (to figure out which features to add)? Or maybe do hierarchical clustering?",
          "votes": 3
        }
      ]
    },
    {
      "id": 1878709,
      "postDate": "2022-07-31T15:04:01.573Z",
      "content": "<p>How do you measure predictiveness ? it seems that some important categoricals are missing<br>\nHow do you match the target with the data ? keeping only last row for customer ? </p>\n<p>Edit: why the downvote ?</p>",
      "rawMarkdown": "How do you measure predictiveness ? it seems that some important categoricals are missing\nHow do you match the target with the data ? keeping only last row for customer ? \n\nEdit: why the downvote ?",
      "votes": 2,
      "replies": [
        {
          "id": 1878715,
          "postDate": "2022-07-31T15:10:15.237Z",
          "content": "<p>Ah, I just did the most basic thing I could imagine, which was to just run catboost on each columns and try to have it predict the target. </p>\n<p>That might have been a bad idea and maybe I should have used a correlation matrix?</p>\n<p>To match the target with the data, I just did a majority vote on my test row predictions, so I just took the mode to find out what label occurred most often for that customer (0 or 1; in the event of a tie, I picked 0).</p>\n<p>Just so you know, I didn't downvote you! : )</p>",
          "rawMarkdown": "Ah, I just did the most basic thing I could imagine, which was to just run catboost on each columns and try to have it predict the target. \n\nThat might have been a bad idea and maybe I should have used a correlation matrix?\n\nTo match the target with the data, I just did a majority vote on my test row predictions, so I just took the mode to find out what label occurred most often for that customer (0 or 1; in the event of a tie, I picked 0).\n\nJust so you know, I didn't downvote you! : )",
          "votes": 1
        },
        {
          "id": 1878724,
          "postDate": "2022-07-31T15:24:36.543Z",
          "content": "<p>Regarding features, you are probably missing some interactions. Usually not a good idea to remove features in an univariate fashion… The other problem seems to come from the way you match data and target. We have multiple row for each clients but only one target. It seems you have duplicated the target on each rows of the clients… and similarly you predict multiple rows then agregate them. Try a model on the last row of each client to predict the target. It will give better results.</p>",
          "rawMarkdown": "Regarding features, you are probably missing some interactions. Usually not a good idea to remove features in an univariate fashion... The other problem seems to come from the way you match data and target. We have multiple row for each clients but only one target. It seems you have duplicated the target on each rows of the clients... and similarly you predict multiple rows then agregate them. Try a model on the last row of each client to predict the target. It will give better results.",
          "votes": 3
        },
        {
          "id": 1878742,
          "postDate": "2022-07-31T15:43:38.783Z",
          "content": "<p>I agree with this, use the last row for the client and use the target predictions thereby. Also, focus on feature engineering and perhaps some better feature selection methods. <br>\nGood luck!</p>",
          "rawMarkdown": "I agree with this, use the last row for the client and use the target predictions thereby. Also, focus on feature engineering and perhaps some better feature selection methods. \nGood luck!"
        },
        {
          "id": 1878846,
          "postDate": "2022-07-31T17:06:27.713Z",
          "content": "<p>Thank you folks!!! That sounds like good advice!</p>",
          "rawMarkdown": "Thank you folks!!! That sounds like good advice!"
        }
      ]
    },
    {
      "id": 1878720,
      "postDate": "2022-07-31T15:14:44.680Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1878860,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2022-07-31T17:16:05.440000",
      "content": "<p>I'd suggest these next steps:</p>\n<ol>\n<li>As <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> recommends: Keep only the last row for every customer (i.e. the most recent data) and get rid of the majority vote.</li>\n<li>Don't submit the binary labels, but the predicted probabilities (CatBoostClassifier.predict_proba). The competition's evaluation metric needs the probabilities. (You'd submit the labels if the competition were evaluated by accuracy.)</li>\n<li>Don't be afraid of overfitting: Use large trees and unlimited iterations with early stopping.</li>\n<li>Add some more features.</li>\n<li>Play with the catboost parameter colsample_bylevel. You don't need Optuna for that.</li>\n</ol>\n<p>And tell us about the results…</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1880105,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-08-01T13:17:23.977000",
          "content": "<p>I still struggle with 3). Would you recommend this outside of kaggle ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1880125,
          "author_name": "Campbell Hutcheson",
          "author_url": "",
          "post_date": "2022-08-01T13:30:45.217000",
          "content": "<p>Great advice! </p>\n<p>I was able to get to .741 : O using just 1) and 2), so now I'll take a look at 3), 4) and 5)!</p>\n<p>I am using InterpretML to try to get an idea of feature interactions, I don't see anyone else using it on here, though.</p>\n<p>Do folks here normally look at t-SNE plots to try to get an idea of feature clusters (to figure out which features to add)? Or maybe do hierarchical clustering?</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1878709,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-07-31T15:04:01.573000",
      "content": "<p>How do you measure predictiveness ? it seems that some important categoricals are missing<br>\nHow do you match the target with the data ? keeping only last row for customer ? </p>\n<p>Edit: why the downvote ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1878715,
          "author_name": "Campbell Hutcheson",
          "author_url": "",
          "post_date": "2022-07-31T15:10:15.237000",
          "content": "<p>Ah, I just did the most basic thing I could imagine, which was to just run catboost on each columns and try to have it predict the target. </p>\n<p>That might have been a bad idea and maybe I should have used a correlation matrix?</p>\n<p>To match the target with the data, I just did a majority vote on my test row predictions, so I just took the mode to find out what label occurred most often for that customer (0 or 1; in the event of a tie, I picked 0).</p>\n<p>Just so you know, I didn't downvote you! : )</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1878724,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-07-31T15:24:36.543000",
          "content": "<p>Regarding features, you are probably missing some interactions. Usually not a good idea to remove features in an univariate fashion… The other problem seems to come from the way you match data and target. We have multiple row for each clients but only one target. It seems you have duplicated the target on each rows of the clients… and similarly you predict multiple rows then agregate them. Try a model on the last row of each client to predict the target. It will give better results.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1878742,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2022-07-31T15:43:38.783000",
          "content": "<p>I agree with this, use the last row for the client and use the target predictions thereby. Also, focus on feature engineering and perhaps some better feature selection methods. <br>\nGood luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1878846,
          "author_name": "Campbell Hutcheson",
          "author_url": "",
          "post_date": "2022-07-31T17:06:27.713000",
          "content": "<p>Thank you folks!!! That sounds like good advice!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1878720,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-31T15:14:44.680000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1878670": "So, I'm trying to get to .6 without using any of the publicly available code as a practice. I'm at .451 and wondered if anyone might have any advice.\n\n1) After creating CV folds, I cut down the dataset to just the following columns to make my models faster to train and the data easier to explore.\n\nI picked the following columns just based upon single column predictiveness of the target.\n\n[\"customer_ID\", \"P_2\", \"B_18\", \"B_9\", \"B_2\", \"D_48\", \"D_44\", \"B_3\", \"B_1\", \"B_37\", \"D_75\", \"B_11\", \"B_19\", \"B_22\", \"D_61\", \"B_6\", \"B_7\", \"B_23\", \"D_62\", \"B_38\", \"D_55\"]\n\nMy intuition is that the number of columns that actually matter must be very small and most of the columns must be mostly irrelevant.\n\n2) I threw the data into Catboost as a first attempt and used the scale_pos_weight with a 4:1 weighting to reflect the imbalance of the data. \n\nI limited the model to a small number of iterations and small tree size because I expect that there is a real risk of overfitting.\n\n3) I did predictions on the test data rows and then did a majority vote per customer ID to reach a prediction.\n\nWhat do you think would be the best next steps?\n\n1) Is it possible to hit .6 with just the features that I picked or should my next step be to add more features?\n\n2) I could just a hyper parameter optimization tool like Optuna or Skopt but I feel that with my results as low as they are, there is probably something fundamentally wrong that I need to fix first.",
    "1878860": "I'd suggest these next steps:\n1. As @lucasmorin recommends: Keep only the last row for every customer (i.e. the most recent data) and get rid of the majority vote.\n2. Don't submit the binary labels, but the predicted probabilities (CatBoostClassifier.predict_proba). The competition's evaluation metric needs the probabilities. (You'd submit the labels if the competition were evaluated by accuracy.)\n3. Don't be afraid of overfitting: Use large trees and unlimited iterations with early stopping.\n4. Add some more features.\n5. Play with the catboost parameter colsample_bylevel. You don't need Optuna for that.\n\nAnd tell us about the results...",
    "1878709": "How do you measure predictiveness ? it seems that some important categoricals are missing\nHow do you match the target with the data ? keeping only last row for customer ? \n\nEdit: why the downvote ?",
    "1878720": ""
  }
}