{
  "id": 339278,
  "title": "How good are the predictions? See for yourself!",
  "url": "/competitions/amex-default-prediction/discussion/339278",
  "author_name": "",
  "post_date": "2022-07-24T03:28:46.005845700Z",
  "votes": 43,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I like to see how the predictions look like beyond what we get from a 1D array of numbers. So here is my attempt to help us understand how good our predictions are, and why there are problems.</p>\n<p>To set the stage: the points shown below are from one validation fold, so 91700 or so. The predictions were made by a Keras network (LSTM, followed by 64- and 32-unit dense layers). The activations from that last dense layer are fed into a sigmoid layer which makes decisions. For the sake of completeness, I will share that the overall prediction has a CV=0.7895 and LB=0.790. Yet here we are taking those penultimate dense layer activations (the networks has basically learned everything at that point and only a decision remains) and feeding them into <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>t-SNE</strong></a> to get a 2D embedding. Now, what t-SNE does with these 32-dimensional vectors is not the same as what the sigmoid layer does, but in general data points that are well-separated by t-SNE will also be predicted well clear of each other by the network.</p>\n<p>By the way, to see this image properly, you will probably need to right-hand click while hovering over it and do <code>Open image in new tab</code>.  In that new tab you will get a magnifying glass and should be able to zoom in and see the details.</p>\n<p><img src=\"https://i.ibb.co/rcjSyTs/t-SNE-keras-training-02.png\" alt=\"t-SNE plot\"></p>\n<p>A couple of things should be obvious. The neural network has learned fairly well to separate the two groups (and boosted trees have likely done so even better). If you look to the upper left corner at 10:30 in wall clock time, there are several blue dots peppering the island of red dots. Those are defaulted customers we will never be able to predict correctly, no matter the feature engineering or other fancy techniques. That's because they look like non-default customers, and probably something dramatic happens in their lives, something that can't be accounted for or predicted, that makes them default. A similar argument goes for the sparse red dots in two tiny blue islands at 6 o'clock, except that these are customers who look like they should be defaulting, yet somehow they don't. Let's not worry about either group because they are relatively rare, and besides we probably can't do anything about them.</p>\n<p>There are two groups of most likely targets for the improved predictions. One is on the boundary, which I will tentatively call the zone of confusion. Specifically, in that strip between <code>[0, -100]</code> on the axis 2 (Y-axis, if you will). In probability parlance, those are points with approximate values in the <code>[0.35, 0.65]</code> range, and they may be on the right or wrong side of the boundary. This is where the feature engineering helps the most, because it is possible to come up with clever features that will push them to the correct side and out of the zone of confusion. That will definitely improve the accuracy and log-loss of our models. Unfortunately, it is not very likely to improve the AmEx score, as most of those data points don't have high enough probability to be anywhere near the extremes of predictions.</p>\n<p>This brings us to the second category for possible improvement. Those are red dots that are peppering the blue side outside of the zone of confusion, and <em>vice versa</em>. Most of them are not as hopeless as those blue dots at 10:30, but they are still deep in enemy territory. They are most likely the customers AmEx wants us to get right, and is targeting them with the scoring scheme that has brought us all so much joy. Keep in mind that we don't have to pull these predictions all the way from the dark side. It should be enough to bring them into the zone of confusion, or even close to it.</p>\n<p><strong>EDIT</strong>: There is <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451\" target=\"_blank\"><strong>a companion post</strong></a> showing similar type of analysis for LightGBM predictions.</p>\n<p><strong>EDIT #2</strong>: See the conclusion of this topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726\" target=\"_blank\"><strong>here</strong></a>.</p>",
  "messages": [
    {
      "id": "1868527",
      "postDate": "07/24/2022 03:28:46",
      "content": "<p>I like to see how the predictions look like beyond what we get from a 1D array of numbers. So here is my attempt to help us understand how good our predictions are, and why there are problems.</p>\n<p>To set the stage: the points shown below are from one validation fold, so 91700 or so. The predictions were made by a Keras network (LSTM, followed by 64- and 32-unit dense layers). The activations from that last dense layer are fed into a sigmoid layer which makes decisions. For the sake of completeness, I will share that the overall prediction has a CV=0.7895 and LB=0.790. Yet here we are taking those penultimate dense layer activations (the networks has basically learned everything at that point and only a decision remains) and feeding them into <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>t-SNE</strong></a> to get a 2D embedding. Now, what t-SNE does with these 32-dimensional vectors is not the same as what the sigmoid layer does, but in general data points that are well-separated by t-SNE will also be predicted well clear of each other by the network.</p>\n<p>By the way, to see this image properly, you will probably need to right-hand click while hovering over it and do <code>Open image in new tab</code>.  In that new tab you will get a magnifying glass and should be able to zoom in and see the details.</p>\n<p><img src=\"https://i.ibb.co/rcjSyTs/t-SNE-keras-training-02.png\" alt=\"t-SNE plot\"></p>\n<p>A couple of things should be obvious. The neural network has learned fairly well to separate the two groups (and boosted trees have likely done so even better). If you look to the upper left corner at 10:30 in wall clock time, there are several blue dots peppering the island of red dots. Those are defaulted customers we will never be able to predict correctly, no matter the feature engineering or other fancy techniques. That's because they look like non-default customers, and probably something dramatic happens in their lives, something that can't be accounted for or predicted, that makes them default. A similar argument goes for the sparse red dots in two tiny blue islands at 6 o'clock, except that these are customers who look like they should be defaulting, yet somehow they don't. Let's not worry about either group because they are relatively rare, and besides we probably can't do anything about them.</p>\n<p>There are two groups of most likely targets for the improved predictions. One is on the boundary, which I will tentatively call the zone of confusion. Specifically, in that strip between <code>[0, -100]</code> on the axis 2 (Y-axis, if you will). In probability parlance, those are points with approximate values in the <code>[0.35, 0.65]</code> range, and they may be on the right or wrong side of the boundary. This is where the feature engineering helps the most, because it is possible to come up with clever features that will push them to the correct side and out of the zone of confusion. That will definitely improve the accuracy and log-loss of our models. Unfortunately, it is not very likely to improve the AmEx score, as most of those data points don't have high enough probability to be anywhere near the extremes of predictions.</p>\n<p>This brings us to the second category for possible improvement. Those are red dots that are peppering the blue side outside of the zone of confusion, and <em>vice versa</em>. Most of them are not as hopeless as those blue dots at 10:30, but they are still deep in enemy territory. They are most likely the customers AmEx wants us to get right, and is targeting them with the scoring scheme that has brought us all so much joy. Keep in mind that we don't have to pull these predictions all the way from the dark side. It should be enough to bring them into the zone of confusion, or even close to it.</p>\n<p><strong>EDIT</strong>: There is <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451\" target=\"_blank\"><strong>a companion post</strong></a> showing similar type of analysis for LightGBM predictions.</p>\n<p><strong>EDIT #2</strong>: See the conclusion of this topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726\" target=\"_blank\"><strong>here</strong></a>.</p>",
      "rawMarkdown": "I like to see how the predictions look like beyond what we get from a 1D array of numbers. So here is my attempt to help us understand how good our predictions are, and why there are problems.\n\nTo set the stage: the points shown below are from one validation fold, so 91700 or so. The predictions were made by a Keras network (LSTM, followed by 64- and 32-unit dense layers). The activations from that last dense layer are fed into a sigmoid layer which makes decisions. For the sake of completeness, I will share that the overall prediction has a CV=0.7895 and LB=0.790. Yet here we are taking those penultimate dense layer activations (the networks has basically learned everything at that point and only a decision remains) and feeding them into [**t-SNE**](https://github.com/pavlin-policar/openTSNE) to get a 2D embedding. Now, what t-SNE does with these 32-dimensional vectors is not the same as what the sigmoid layer does, but in general data points that are well-separated by t-SNE will also be predicted well clear of each other by the network.\n\nBy the way, to see this image properly, you will probably need to right-hand click while hovering over it and do `Open image in new tab`.  In that new tab you will get a magnifying glass and should be able to zoom in and see the details.\n\n![t-SNE plot](https://i.ibb.co/rcjSyTs/t-SNE-keras-training-02.png)\n\nA couple of things should be obvious. The neural network has learned fairly well to separate the two groups (and boosted trees have likely done so even better). If you look to the upper left corner at 10:30 in wall clock time, there are several blue dots peppering the island of red dots. Those are defaulted customers we will never be able to predict correctly, no matter the feature engineering or other fancy techniques. That's because they look like non-default customers, and probably something dramatic happens in their lives, something that can't be accounted for or predicted, that makes them default. A similar argument goes for the sparse red dots in two tiny blue islands at 6 o'clock, except that these are customers who look like they should be defaulting, yet somehow they don't. Let's not worry about either group because they are relatively rare, and besides we probably can't do anything about them.\n\nThere are two groups of most likely targets for the improved predictions. One is on the boundary, which I will tentatively call the zone of confusion. Specifically, in that strip between `[0, -100]` on the axis 2 (Y-axis, if you will). In probability parlance, those are points with approximate values in the `[0.35, 0.65]` range, and they may be on the right or wrong side of the boundary. This is where the feature engineering helps the most, because it is possible to come up with clever features that will push them to the correct side and out of the zone of confusion. That will definitely improve the accuracy and log-loss of our models. Unfortunately, it is not very likely to improve the AmEx score, as most of those data points don't have high enough probability to be anywhere near the extremes of predictions.\n\nThis brings us to the second category for possible improvement. Those are red dots that are peppering the blue side outside of the zone of confusion, and *vice versa*. Most of them are not as hopeless as those blue dots at 10:30, but they are still deep in enemy territory. They are most likely the customers AmEx wants us to get right, and is targeting them with the scoring scheme that has brought us all so much joy. Keep in mind that we don't have to pull these predictions all the way from the dark side. It should be enough to bring them into the zone of confusion, or even close to it.\n\n**EDIT**: There is [**a companion post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451) showing similar type of analysis for LightGBM predictions.\n\n**EDIT #2**: See the conclusion of this topic [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726).",
      "votes": null
    },
    {
      "id": "1868703",
      "postDate": "07/24/2022 06:56:22",
      "content": "<p>Very well written Sir! Your posts are highly useful in general. Thanks for providing such useful content!</p>",
      "rawMarkdown": "Very well written Sir! Your posts are highly useful in general. Thanks for providing such useful content!",
      "votes": null
    },
    {
      "id": "1868763",
      "postDate": "07/24/2022 07:47:06",
      "content": "<p>Glad you liked it.</p>",
      "rawMarkdown": "Glad you liked it.",
      "votes": null
    },
    {
      "id": "1869218",
      "postDate": "07/24/2022 15:27:18",
      "content": "<p>how can we make similar plot analysis if we are using boosted tree ? </p>",
      "rawMarkdown": "how can we make similar plot analysis if we are using boosted tree ?",
      "votes": null
    },
    {
      "id": "1869223",
      "postDate": "07/24/2022 15:29:58",
      "content": "<p>To get those scores from an LSTM model is impressive. Well done! </p>",
      "rawMarkdown": "To get those scores from an LSTM model is impressive. Well done!",
      "votes": null
    },
    {
      "id": "1869268",
      "postDate": "07/24/2022 16:09:58",
      "content": "<p>Data is beautiful! </p>",
      "rawMarkdown": "Data is beautiful!",
      "votes": null
    },
    {
      "id": "1869298",
      "postDate": "07/24/2022 16:48:22",
      "content": "<p>Still hoping for a better score, because these models ensemble really well with boosted trees.</p>",
      "rawMarkdown": "Still hoping for a better score, because these models ensemble really well with boosted trees.",
      "votes": null
    },
    {
      "id": "1869301",
      "postDate": "07/24/2022 16:49:40",
      "content": "<p>A similar plot can be done from leaf indices of boosted trees. I will make it one of these days.</p>\n<p><strong>EDIT</strong>: I made a similar plot based on LightGBM in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451\" target=\"_blank\"><strong>this post</strong></a>.</p>",
      "rawMarkdown": "A similar plot can be done from leaf indices of boosted trees. I will make it one of these days.\n\n**EDIT**: I made a similar plot based on LightGBM in [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451).",
      "votes": null
    },
    {
      "id": "1870040",
      "postDate": "07/25/2022 08:17:13",
      "content": "<p>This is amazing! I will try to see this whenever I'll use a NN classifier now ;-)</p>",
      "rawMarkdown": "This is amazing! I will try to see this whenever I'll use a NN classifier now ;-)",
      "votes": null
    },
    {
      "id": "1871789",
      "postDate": "07/26/2022 13:33:22",
      "content": "<p>Fascinating exploration, thanks for sharing! Do you know roughly how t-SNE is reducing this data down? </p>",
      "rawMarkdown": "Fascinating exploration, thanks for sharing! Do you know roughly how t-SNE is reducing this data down?",
      "votes": null
    },
    {
      "id": "1872271",
      "postDate": "07/26/2022 19:00:33",
      "content": "<p>It is known how t-SNE works. It is a non-linear dimensionality reduction method that preserves the neighborhood of data points, such that similar samples end up next to each other. This is usually helpful for visual inspection, but it doesn't preserve global distances between samples. If you look at <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726#1871426\" target=\"_blank\"><strong>this plot</strong></a> where data points are colored by what neural network will eventually predict, there is a nice color continuity from one tip of the worm to the other, meaning that locally everything looks right. Yet the blue tip is closer to the middle than the yellow tip, and they should be equidistant. Also, the two tips are too close to each other, yet they should be as far away as possible based on their global separation in raw data.</p>\n<p>PCA, for example, is a linear dimensionality reduction, and all distances in a 2D plot are proportional to actual sample distances. Even though in mathematical sense PCA plots are more appropriate (as you will see in a second), they are often difficult to visualize, especially on complex datasets. On a \"simple\" dataset such as the one referenced in my link above, here is what PCA will produce:</p>\n<p><img src=\"https://i.ibb.co/BcgxbPw/PCA-keras-training-06.png\" alt=\"PCA plot\"></p>\n<p>In this case the tips are as far away as possible, and they are roughly equidistant to the middle, but keep in mind again that this is a relatively straightforward dataset. Still, why are we using t-SNE when in this case PCA does a better job? Look no further than the image on the top of this page, and compare it to the PCA transformation on the same dataset which is shown below.</p>\n<p><img src=\"https://i.ibb.co/DKDFgZ7/PCA-keras-training-02.png\" alt=\"PCA plot\"></p>\n<p>One should never assume that all of us see things the same way, but I think it is fair to say that most people would find t-SNE rendering from above to be more amenable to interpretation. For complex datasets it is often easier to see the underlying data structure the way t-SNE puts it into the plot, but we must keep in mind that global distances between points are not preserved.</p>",
      "rawMarkdown": "It is known how t-SNE works. It is a non-linear dimensionality reduction method that preserves the neighborhood of data points, such that similar samples end up next to each other. This is usually helpful for visual inspection, but it doesn't preserve global distances between samples. If you look at [**this plot**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726#1871426) where data points are colored by what neural network will eventually predict, there is a nice color continuity from one tip of the worm to the other, meaning that locally everything looks right. Yet the blue tip is closer to the middle than the yellow tip, and they should be equidistant. Also, the two tips are too close to each other, yet they should be as far away as possible based on their global separation in raw data.\n\nPCA, for example, is a linear dimensionality reduction, and all distances in a 2D plot are proportional to actual sample distances. Even though in mathematical sense PCA plots are more appropriate (as you will see in a second), they are often difficult to visualize, especially on complex datasets. On a \"simple\" dataset such as the one referenced in my link above, here is what PCA will produce:\n\n![PCA plot](https://i.ibb.co/BcgxbPw/PCA-keras-training-06.png)\n\nIn this case the tips are as far away as possible, and they are roughly equidistant to the middle, but keep in mind again that this is a relatively straightforward dataset. Still, why are we using t-SNE when in this case PCA does a better job? Look no further than the image on the top of this page, and compare it to the PCA transformation on the same dataset which is shown below.\n\n![PCA plot](https://i.ibb.co/DKDFgZ7/PCA-keras-training-02.png)\n\nOne should never assume that all of us see things the same way, but I think it is fair to say that most people would find t-SNE rendering from above to be more amenable to interpretation. For complex datasets it is often easier to see the underlying data structure the way t-SNE puts it into the plot, but we must keep in mind that global distances between points are not preserved.",
      "votes": null
    },
    {
      "id": "1872386",
      "postDate": "07/26/2022 22:25:41",
      "content": "<p>Great visualization of prediction results! I really like the insights that you shared.</p>",
      "rawMarkdown": "Great visualization of prediction results! I really like the insights that you shared.",
      "votes": null
    },
    {
      "id": "1873075",
      "postDate": "07/27/2022 12:36:50",
      "content": "<p>The way t-sne has separated the manifold of learning dexterity between default and non default is awesome. </p>",
      "rawMarkdown": "The way t-sne has separated the manifold of learning dexterity between default and non default is awesome.",
      "votes": null
    },
    {
      "id": "1873418",
      "postDate": "07/27/2022 15:43:04",
      "content": "<p>The separation gets better from stacked predictions, which is shown <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726\" target=\"_blank\"><strong>here</strong></a>.</p>",
      "rawMarkdown": "The separation gets better from stacked predictions, which is shown [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726).",
      "votes": null
    },
    {
      "id": "1873433",
      "postDate": "07/27/2022 15:52:56",
      "content": "<p>thanks sir for this</p>",
      "rawMarkdown": "thanks sir for this",
      "votes": null
    },
    {
      "id": "1882125",
      "postDate": "08/03/2022 04:59:37",
      "content": "<p>This is amazing! Great job</p>",
      "rawMarkdown": "This is amazing! Great job",
      "votes": null
    },
    {
      "id": "1882138",
      "postDate": "08/03/2022 05:08:52",
      "content": "<p>Hey!! This is great, you are great teaching!</p>",
      "rawMarkdown": "Hey!! This is great, you are great teaching!",
      "votes": null
    },
    {
      "id": "1882141",
      "postDate": "08/03/2022 05:09:41",
      "content": "<p>I really like your post, you are really good with your explanation. Amazing plot!</p>",
      "rawMarkdown": "I really like your post, you are really good with your explanation. Amazing plot!",
      "votes": null
    },
    {
      "id": "1883622",
      "postDate": "08/04/2022 01:15:09",
      "content": "<p>Thanks for your work! It's very helpful.</p>",
      "rawMarkdown": "Thanks for your work! It's very helpful.",
      "votes": null
    },
    {
      "id": "1883686",
      "postDate": "08/04/2022 03:05:24",
      "content": "<p>I have noticed that dreaded horseshoe pattern in many a PCA visualization. I definitely appreciate what t-SNE (or UMAP) offer.</p>",
      "rawMarkdown": "I have noticed that dreaded horseshoe pattern in many a PCA visualization. I definitely appreciate what t-SNE (or UMAP) offer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1868703,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/24/2022 06:56:22",
      "content": "<p>Very well written Sir! Your posts are highly useful in general. Thanks for providing such useful content!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1868763,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/24/2022 07:47:06",
          "content": "<p>Glad you liked it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869218,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "07/24/2022 15:27:18",
      "content": "<p>how can we make similar plot analysis if we are using boosted tree ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1869301,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/24/2022 16:49:40",
          "content": "<p>A similar plot can be done from leaf indices of boosted trees. I will make it one of these days.</p>\n<p><strong>EDIT</strong>: I made a similar plot based on LightGBM in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451\" target=\"_blank\"><strong>this post</strong></a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869223,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/24/2022 15:29:58",
      "content": "<p>To get those scores from an LSTM model is impressive. Well done! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1869298,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/24/2022 16:48:22",
          "content": "<p>Still hoping for a better score, because these models ensemble really well with boosted trees.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869268,
      "author_name": "jakelj",
      "author_url": "",
      "post_date": "07/24/2022 16:09:58",
      "content": "<p>Data is beautiful! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1870040,
      "author_name": "virajkadam",
      "author_url": "",
      "post_date": "07/25/2022 08:17:13",
      "content": "<p>This is amazing! I will try to see this whenever I'll use a NN classifier now ;-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1871789,
      "author_name": "codeslang",
      "author_url": "",
      "post_date": "07/26/2022 13:33:22",
      "content": "<p>Fascinating exploration, thanks for sharing! Do you know roughly how t-SNE is reducing this data down? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1872271,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/26/2022 19:00:33",
          "content": "<p>It is known how t-SNE works. It is a non-linear dimensionality reduction method that preserves the neighborhood of data points, such that similar samples end up next to each other. This is usually helpful for visual inspection, but it doesn't preserve global distances between samples. If you look at <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726#1871426\" target=\"_blank\"><strong>this plot</strong></a> where data points are colored by what neural network will eventually predict, there is a nice color continuity from one tip of the worm to the other, meaning that locally everything looks right. Yet the blue tip is closer to the middle than the yellow tip, and they should be equidistant. Also, the two tips are too close to each other, yet they should be as far away as possible based on their global separation in raw data.</p>\n<p>PCA, for example, is a linear dimensionality reduction, and all distances in a 2D plot are proportional to actual sample distances. Even though in mathematical sense PCA plots are more appropriate (as you will see in a second), they are often difficult to visualize, especially on complex datasets. On a \"simple\" dataset such as the one referenced in my link above, here is what PCA will produce:</p>\n<p><img src=\"https://i.ibb.co/BcgxbPw/PCA-keras-training-06.png\" alt=\"PCA plot\"></p>\n<p>In this case the tips are as far away as possible, and they are roughly equidistant to the middle, but keep in mind again that this is a relatively straightforward dataset. Still, why are we using t-SNE when in this case PCA does a better job? Look no further than the image on the top of this page, and compare it to the PCA transformation on the same dataset which is shown below.</p>\n<p><img src=\"https://i.ibb.co/DKDFgZ7/PCA-keras-training-02.png\" alt=\"PCA plot\"></p>\n<p>One should never assume that all of us see things the same way, but I think it is fair to say that most people would find t-SNE rendering from above to be more amenable to interpretation. For complex datasets it is often easier to see the underlying data structure the way t-SNE puts it into the plot, but we must keep in mind that global distances between points are not preserved.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1882138,
          "author_name": "agustin222",
          "author_url": "",
          "post_date": "08/03/2022 05:08:52",
          "content": "<p>Hey!! This is great, you are great teaching!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1883686,
          "author_name": "scharlesworth",
          "author_url": "",
          "post_date": "08/04/2022 03:05:24",
          "content": "<p>I have noticed that dreaded horseshoe pattern in many a PCA visualization. I definitely appreciate what t-SNE (or UMAP) offer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1872386,
      "author_name": "oscarm524",
      "author_url": "",
      "post_date": "07/26/2022 22:25:41",
      "content": "<p>Great visualization of prediction results! I really like the insights that you shared.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1873075,
      "author_name": "xtremboot22",
      "author_url": "",
      "post_date": "07/27/2022 12:36:50",
      "content": "<p>The way t-sne has separated the manifold of learning dexterity between default and non default is awesome. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1873418,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/27/2022 15:43:04",
          "content": "<p>The separation gets better from stacked predictions, which is shown <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726\" target=\"_blank\"><strong>here</strong></a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1873433,
      "author_name": "bilalsuppal",
      "author_url": "",
      "post_date": "07/27/2022 15:52:56",
      "content": "<p>thanks sir for this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1882125,
      "author_name": "andrew531",
      "author_url": "",
      "post_date": "08/03/2022 04:59:37",
      "content": "<p>This is amazing! Great job</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1882141,
      "author_name": "agustin222",
      "author_url": "",
      "post_date": "08/03/2022 05:09:41",
      "content": "<p>I really like your post, you are really good with your explanation. Amazing plot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1883622,
      "author_name": "milkyio",
      "author_url": "",
      "post_date": "08/04/2022 01:15:09",
      "content": "<p>Thanks for your work! It's very helpful.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1868527": "I like to see how the predictions look like beyond what we get from a 1D array of numbers. So here is my attempt to help us understand how good our predictions are, and why there are problems.\n\nTo set the stage: the points shown below are from one validation fold, so 91700 or so. The predictions were made by a Keras network (LSTM, followed by 64- and 32-unit dense layers). The activations from that last dense layer are fed into a sigmoid layer which makes decisions. For the sake of completeness, I will share that the overall prediction has a CV=0.7895 and LB=0.790. Yet here we are taking those penultimate dense layer activations (the networks has basically learned everything at that point and only a decision remains) and feeding them into [**t-SNE**](https://github.com/pavlin-policar/openTSNE) to get a 2D embedding. Now, what t-SNE does with these 32-dimensional vectors is not the same as what the sigmoid layer does, but in general data points that are well-separated by t-SNE will also be predicted well clear of each other by the network.\n\nBy the way, to see this image properly, you will probably need to right-hand click while hovering over it and do `Open image in new tab`.  In that new tab you will get a magnifying glass and should be able to zoom in and see the details.\n\n![t-SNE plot](https://i.ibb.co/rcjSyTs/t-SNE-keras-training-02.png)\n\nA couple of things should be obvious. The neural network has learned fairly well to separate the two groups (and boosted trees have likely done so even better). If you look to the upper left corner at 10:30 in wall clock time, there are several blue dots peppering the island of red dots. Those are defaulted customers we will never be able to predict correctly, no matter the feature engineering or other fancy techniques. That's because they look like non-default customers, and probably something dramatic happens in their lives, something that can't be accounted for or predicted, that makes them default. A similar argument goes for the sparse red dots in two tiny blue islands at 6 o'clock, except that these are customers who look like they should be defaulting, yet somehow they don't. Let's not worry about either group because they are relatively rare, and besides we probably can't do anything about them.\n\nThere are two groups of most likely targets for the improved predictions. One is on the boundary, which I will tentatively call the zone of confusion. Specifically, in that strip between `[0, -100]` on the axis 2 (Y-axis, if you will). In probability parlance, those are points with approximate values in the `[0.35, 0.65]` range, and they may be on the right or wrong side of the boundary. This is where the feature engineering helps the most, because it is possible to come up with clever features that will push them to the correct side and out of the zone of confusion. That will definitely improve the accuracy and log-loss of our models. Unfortunately, it is not very likely to improve the AmEx score, as most of those data points don't have high enough probability to be anywhere near the extremes of predictions.\n\nThis brings us to the second category for possible improvement. Those are red dots that are peppering the blue side outside of the zone of confusion, and *vice versa*. Most of them are not as hopeless as those blue dots at 10:30, but they are still deep in enemy territory. They are most likely the customers AmEx wants us to get right, and is targeting them with the scoring scheme that has brought us all so much joy. Keep in mind that we don't have to pull these predictions all the way from the dark side. It should be enough to bring them into the zone of confusion, or even close to it.\n\n**EDIT**: There is [**a companion post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451) showing similar type of analysis for LightGBM predictions.\n\n**EDIT #2**: See the conclusion of this topic [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726).",
    "1868703": "Very well written Sir! Your posts are highly useful in general. Thanks for providing such useful content!",
    "1868763": "Glad you liked it.",
    "1869218": "how can we make similar plot analysis if we are using boosted tree ?",
    "1869223": "To get those scores from an LSTM model is impressive. Well done!",
    "1869268": "Data is beautiful!",
    "1869298": "Still hoping for a better score, because these models ensemble really well with boosted trees.",
    "1869301": "A similar plot can be done from leaf indices of boosted trees. I will make it one of these days.\n\n**EDIT**: I made a similar plot based on LightGBM in [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339451).",
    "1870040": "This is amazing! I will try to see this whenever I'll use a NN classifier now ;-)",
    "1871789": "Fascinating exploration, thanks for sharing! Do you know roughly how t-SNE is reducing this data down?",
    "1872271": "It is known how t-SNE works. It is a non-linear dimensionality reduction method that preserves the neighborhood of data points, such that similar samples end up next to each other. This is usually helpful for visual inspection, but it doesn't preserve global distances between samples. If you look at [**this plot**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726#1871426) where data points are colored by what neural network will eventually predict, there is a nice color continuity from one tip of the worm to the other, meaning that locally everything looks right. Yet the blue tip is closer to the middle than the yellow tip, and they should be equidistant. Also, the two tips are too close to each other, yet they should be as far away as possible based on their global separation in raw data.\n\nPCA, for example, is a linear dimensionality reduction, and all distances in a 2D plot are proportional to actual sample distances. Even though in mathematical sense PCA plots are more appropriate (as you will see in a second), they are often difficult to visualize, especially on complex datasets. On a \"simple\" dataset such as the one referenced in my link above, here is what PCA will produce:\n\n![PCA plot](https://i.ibb.co/BcgxbPw/PCA-keras-training-06.png)\n\nIn this case the tips are as far away as possible, and they are roughly equidistant to the middle, but keep in mind again that this is a relatively straightforward dataset. Still, why are we using t-SNE when in this case PCA does a better job? Look no further than the image on the top of this page, and compare it to the PCA transformation on the same dataset which is shown below.\n\n![PCA plot](https://i.ibb.co/DKDFgZ7/PCA-keras-training-02.png)\n\nOne should never assume that all of us see things the same way, but I think it is fair to say that most people would find t-SNE rendering from above to be more amenable to interpretation. For complex datasets it is often easier to see the underlying data structure the way t-SNE puts it into the plot, but we must keep in mind that global distances between points are not preserved.",
    "1872386": "Great visualization of prediction results! I really like the insights that you shared.",
    "1873075": "The way t-sne has separated the manifold of learning dexterity between default and non default is awesome.",
    "1873418": "The separation gets better from stacked predictions, which is shown [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726).",
    "1873433": "thanks sir for this",
    "1882125": "This is amazing! Great job",
    "1882138": "Hey!! This is great, you are great teaching!",
    "1882141": "I really like your post, you are really good with your explanation. Amazing plot!",
    "1883622": "Thanks for your work! It's very helpful.",
    "1883686": "I have noticed that dreaded horseshoe pattern in many a PCA visualization. I definitely appreciate what t-SNE (or UMAP) offer."
  },
  "source": "meta"
}