{
  "id": 336349,
  "title": "Features to detect outliers?",
  "url": "/competitions/amex-default-prediction/discussion/336349",
  "author_name": "",
  "post_date": "2022-07-10T17:54:13.606151800Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I was planning to add a new feature for each numeric feature to tell the model whether a sample for that feature is an outlier or not. So, basically I thought the new feature would have a 1 against samples that are outliers and 0 against samples that aren't. Since we're supposed to detect fraud, I thought it would be useful to let the model know which values are unusual. Do you guys think this would be useful? I would love to hear your thoughts. I'm sorry if the question is stupid. </p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1850757",
      "postDate": "07/10/2022 17:54:13",
      "content": "<p>I was planning to add a new feature for each numeric feature to tell the model whether a sample for that feature is an outlier or not. So, basically I thought the new feature would have a 1 against samples that are outliers and 0 against samples that aren't. Since we're supposed to detect fraud, I thought it would be useful to let the model know which values are unusual. Do you guys think this would be useful? I would love to hear your thoughts. I'm sorry if the question is stupid. </p>\n<p>Thanks.</p>",
      "rawMarkdown": "I was planning to add a new feature for each numeric feature to tell the model whether a sample for that feature is an outlier or not. So, basically I thought the new feature would have a 1 against samples that are outliers and 0 against samples that aren't. Since we're supposed to detect fraud, I thought it would be useful to let the model know which values are unusual. Do you guys think this would be useful? I would love to hear your thoughts. I'm sorry if the question is stupid. \n\nThanks.",
      "votes": null
    },
    {
      "id": "1850778",
      "postDate": "07/10/2022 18:22:54",
      "content": "<p>I feel this is very good idea. </p>",
      "rawMarkdown": "I feel this is very good idea.",
      "votes": null
    },
    {
      "id": "1850781",
      "postDate": "07/10/2022 18:28:00",
      "content": "<p>I'll try it out. Thanks.</p>",
      "rawMarkdown": "I'll try it out. Thanks.",
      "votes": null
    },
    {
      "id": "1851004",
      "postDate": "07/11/2022 01:11:48",
      "content": "<p>This could definitely be useful, although it might not be necessary to create a new feature for each numeric feature - you could also generate a single outlier feature that is based on all of the numeric features.<br>\nOutliers can be tricky to deal with, so it's definitely worth experiment with different approaches to see what works best on your data.<br>\nGood luck!</p>",
      "rawMarkdown": "This could definitely be useful, although it might not be necessary to create a new feature for each numeric feature - you could also generate a single outlier feature that is based on all of the numeric features.\nOutliers can be tricky to deal with, so it's definitely worth experiment with different approaches to see what works best on your data.\nGood luck!",
      "votes": null
    },
    {
      "id": "1851192",
      "postDate": "07/11/2022 04:39:25",
      "content": "<p>Thanks…I was actually planning to calculate Z-score across each feature and have an outlier feature for each individual feature. I don't understand how I could replace the outliers with a single feature? Can you please explain about that?</p>",
      "rawMarkdown": "Thanks...I was actually planning to calculate Z-score across each feature and have an outlier feature for each individual feature. I don't understand how I could replace the outliers with a single feature? Can you please explain about that?",
      "votes": null
    },
    {
      "id": "1851353",
      "postDate": "07/11/2022 07:23:35",
      "content": "<p>That is a great idea. we know that outliers can be problematic for boosting methods (because boosting builds each tree on previous trees' residuals/errors). I  suggest two experiments: </p>\n<ol>\n<li>remove the outliers and run your baseline code with 3 different seed and average CV scores</li>\n<li>add new features stating if there is outlier or not<br>\nI feel the first option will give you better results.</li>\n</ol>",
      "rawMarkdown": "That is a great idea. we know that outliers can be problematic for boosting methods (because boosting builds each tree on previous trees' residuals/errors). I  suggest two experiments: \n1. remove the outliers and run your baseline code with 3 different seed and average CV scores\n2. add new features stating if there is outlier or not\nI feel the first option will give you better results.",
      "votes": null
    },
    {
      "id": "1851369",
      "postDate": "07/11/2022 07:34:40",
      "content": "<p>Assuming that your model is based on decision trees, I doubt that the additional features would be useful. A decision tree detects by itself on what feature and what threshold for this feature it has to branch. A boolean feature for outliers doesn't provide any additional value to the algorithm.</p>\n<p>If you belong to the minority who uses linear models for this competition 😏, the situation would be different: Linear models are very bad at dealing with outliers, and you'd have to apply suitable feature engineering.</p>",
      "rawMarkdown": "Assuming that your model is based on decision trees, I doubt that the additional features would be useful. A decision tree detects by itself on what feature and what threshold for this feature it has to branch. A boolean feature for outliers doesn't provide any additional value to the algorithm.\n\nIf you belong to the minority who uses linear models for this competition 😏, the situation would be different: Linear models are very bad at dealing with outliers, and you'd have to apply suitable feature engineering.",
      "votes": null
    },
    {
      "id": "1851464",
      "postDate": "07/11/2022 09:04:47",
      "content": "<p>Thanks..I'll try both. I'm already trying 2nd actually..After that I'll try first one too.</p>",
      "rawMarkdown": "Thanks..I'll try both. I'm already trying 2nd actually..After that I'll try first one too.",
      "votes": null
    },
    {
      "id": "1851468",
      "postDate": "07/11/2022 09:11:48",
      "content": "<p>Thanks for the tip. I'm certainly not using linear models😄</p>",
      "rawMarkdown": "Thanks for the tip. I'm certainly not using linear models😄",
      "votes": null
    },
    {
      "id": "1851894",
      "postDate": "07/11/2022 15:58:49",
      "content": "<p>Was thinking on similar lines .I assume that one that <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> showcased is basically like a difference from mean <em>X_last - Mean</em> .. which seems to be have some resemblance with Z scores and detecting anomalies that way . I might be wrong but this is the only possible reasoning I thought .</p>\n<p>But as Ambrose commented boosting models are really good at figuring those out as we have seen with some folks generating a lot of em and boosting models finding the best fit .</p>",
      "rawMarkdown": "Was thinking on similar lines .I assume that one that @ragnar123 showcased is basically like a difference from mean *X_last - Mean* .. which seems to be have some resemblance with Z scores and detecting anomalies that way . I might be wrong but this is the only possible reasoning I thought .\n\nBut as Ambrose commented boosting models are really good at figuring those out as we have seen with some folks generating a lot of em and boosting models finding the best fit .",
      "votes": null
    },
    {
      "id": "1851968",
      "postDate": "07/11/2022 17:16:47",
      "content": "<p>I believe you are thinking to implement something similar to : <strong>Calculate Z score. If Z score&gt;3, print it as an outlier</strong><br>\nBut before doing this we need to check if the distribution is normal or not, if normal distribution  then only outlier detection using z-score can be implemented. - others can comment and add upon this.<br>\nNote : not sure if we can apply the same with non-normal distribution, as we have most of the numerical feature in this competition are not normally distributed.</p>\n<p>But AI/ML is all about experimenting , so go for it. All the best.</p>",
      "rawMarkdown": "I believe you are thinking to implement something similar to : **Calculate Z score. If Z score>3, print it as an outlier**\nBut before doing this we need to check if the distribution is normal or not, if normal distribution  then only outlier detection using z-score can be implemented. - others can comment and add upon this.\nNote : not sure if we can apply the same with non-normal distribution, as we have most of the numerical feature in this competition are not normally distributed.\n\nBut AI/ML is all about experimenting , so go for it. All the best.",
      "votes": null
    },
    {
      "id": "1851971",
      "postDate": "07/11/2022 17:17:31",
      "content": "<p>good luck!</p>",
      "rawMarkdown": "good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1850778,
      "author_name": "maheshak04",
      "author_url": "",
      "post_date": "07/10/2022 18:22:54",
      "content": "<p>I feel this is very good idea. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1850781,
          "author_name": "vivekjan14",
          "author_url": "",
          "post_date": "07/10/2022 18:28:00",
          "content": "<p>I'll try it out. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1851004,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/11/2022 01:11:48",
      "content": "<p>This could definitely be useful, although it might not be necessary to create a new feature for each numeric feature - you could also generate a single outlier feature that is based on all of the numeric features.<br>\nOutliers can be tricky to deal with, so it's definitely worth experiment with different approaches to see what works best on your data.<br>\nGood luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1851192,
          "author_name": "vivekjan14",
          "author_url": "",
          "post_date": "07/11/2022 04:39:25",
          "content": "<p>Thanks…I was actually planning to calculate Z-score across each feature and have an outlier feature for each individual feature. I don't understand how I could replace the outliers with a single feature? Can you please explain about that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851968,
          "author_name": "abhianalytic",
          "author_url": "",
          "post_date": "07/11/2022 17:16:47",
          "content": "<p>I believe you are thinking to implement something similar to : <strong>Calculate Z score. If Z score&gt;3, print it as an outlier</strong><br>\nBut before doing this we need to check if the distribution is normal or not, if normal distribution  then only outlier detection using z-score can be implemented. - others can comment and add upon this.<br>\nNote : not sure if we can apply the same with non-normal distribution, as we have most of the numerical feature in this competition are not normally distributed.</p>\n<p>But AI/ML is all about experimenting , so go for it. All the best.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1851353,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/11/2022 07:23:35",
      "content": "<p>That is a great idea. we know that outliers can be problematic for boosting methods (because boosting builds each tree on previous trees' residuals/errors). I  suggest two experiments: </p>\n<ol>\n<li>remove the outliers and run your baseline code with 3 different seed and average CV scores</li>\n<li>add new features stating if there is outlier or not<br>\nI feel the first option will give you better results.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1851464,
          "author_name": "vivekjan14",
          "author_url": "",
          "post_date": "07/11/2022 09:04:47",
          "content": "<p>Thanks..I'll try both. I'm already trying 2nd actually..After that I'll try first one too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1851971,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/11/2022 17:17:31",
          "content": "<p>good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1851369,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "07/11/2022 07:34:40",
      "content": "<p>Assuming that your model is based on decision trees, I doubt that the additional features would be useful. A decision tree detects by itself on what feature and what threshold for this feature it has to branch. A boolean feature for outliers doesn't provide any additional value to the algorithm.</p>\n<p>If you belong to the minority who uses linear models for this competition 😏, the situation would be different: Linear models are very bad at dealing with outliers, and you'd have to apply suitable feature engineering.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1851468,
          "author_name": "vivekjan14",
          "author_url": "",
          "post_date": "07/11/2022 09:11:48",
          "content": "<p>Thanks for the tip. I'm certainly not using linear models😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1851894,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/11/2022 15:58:49",
      "content": "<p>Was thinking on similar lines .I assume that one that <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> showcased is basically like a difference from mean <em>X_last - Mean</em> .. which seems to be have some resemblance with Z scores and detecting anomalies that way . I might be wrong but this is the only possible reasoning I thought .</p>\n<p>But as Ambrose commented boosting models are really good at figuring those out as we have seen with some folks generating a lot of em and boosting models finding the best fit .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1850757": "I was planning to add a new feature for each numeric feature to tell the model whether a sample for that feature is an outlier or not. So, basically I thought the new feature would have a 1 against samples that are outliers and 0 against samples that aren't. Since we're supposed to detect fraud, I thought it would be useful to let the model know which values are unusual. Do you guys think this would be useful? I would love to hear your thoughts. I'm sorry if the question is stupid. \n\nThanks.",
    "1850778": "I feel this is very good idea.",
    "1850781": "I'll try it out. Thanks.",
    "1851004": "This could definitely be useful, although it might not be necessary to create a new feature for each numeric feature - you could also generate a single outlier feature that is based on all of the numeric features.\nOutliers can be tricky to deal with, so it's definitely worth experiment with different approaches to see what works best on your data.\nGood luck!",
    "1851192": "Thanks...I was actually planning to calculate Z-score across each feature and have an outlier feature for each individual feature. I don't understand how I could replace the outliers with a single feature? Can you please explain about that?",
    "1851353": "That is a great idea. we know that outliers can be problematic for boosting methods (because boosting builds each tree on previous trees' residuals/errors). I  suggest two experiments: \n1. remove the outliers and run your baseline code with 3 different seed and average CV scores\n2. add new features stating if there is outlier or not\nI feel the first option will give you better results.",
    "1851369": "Assuming that your model is based on decision trees, I doubt that the additional features would be useful. A decision tree detects by itself on what feature and what threshold for this feature it has to branch. A boolean feature for outliers doesn't provide any additional value to the algorithm.\n\nIf you belong to the minority who uses linear models for this competition 😏, the situation would be different: Linear models are very bad at dealing with outliers, and you'd have to apply suitable feature engineering.",
    "1851464": "Thanks..I'll try both. I'm already trying 2nd actually..After that I'll try first one too.",
    "1851468": "Thanks for the tip. I'm certainly not using linear models😄",
    "1851894": "Was thinking on similar lines .I assume that one that @ragnar123 showcased is basically like a difference from mean *X_last - Mean* .. which seems to be have some resemblance with Z scores and detecting anomalies that way . I might be wrong but this is the only possible reasoning I thought .\n\nBut as Ambrose commented boosting models are really good at figuring those out as we have seen with some folks generating a lot of em and boosting models finding the best fit .",
    "1851968": "I believe you are thinking to implement something similar to : **Calculate Z score. If Z score>3, print it as an outlier**\nBut before doing this we need to check if the distribution is normal or not, if normal distribution  then only outlier detection using z-score can be implemented. - others can comment and add upon this.\nNote : not sure if we can apply the same with non-normal distribution, as we have most of the numerical feature in this competition are not normally distributed.\n\nBut AI/ML is all about experimenting , so go for it. All the best.",
    "1851971": "good luck!"
  },
  "source": "meta"
}