{
  "id": 339726,
  "title": "Better models make better abstract art",
  "url": "/competitions/amex-default-prediction/discussion/339726",
  "author_name": "Tilii",
  "post_date": "2022-07-26T07:55:43.271000",
  "votes": 51,
  "comment_count": 29,
  "views": 0,
  "content": "<p>This is a finishing touch on a topic that started in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>this post</strong></a>. In the original post we had a single neural network model scoring 0.790 on public LB. Here we have 12 different models (a mix of NN, LGB and logistic regression; scores are 0.78-0.795 on public LB) that are stacked together by another neural network. The model scores 0.799 on public LB, and t-SNE of its activations shows much better separation between the two classes. Notice that both tips of the worm have increased density of points, as a stack of models predicts many samples with extreme certainty.</p>\n<p><img src=\"https://i.ibb.co/9rGwK67/t-SNE-keras-training-06.png\" alt=\"t-SNE plot\"></p>\n<p><strong>EDIT</strong>: There is a plot <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271\" target=\"_blank\"><strong>here</strong></a> showing how PCA interprets this dataset.</p>",
  "messages": [
    {
      "id": 1871311,
      "postDate": "2022-07-26T07:55:43.270Z",
      "content": "<p>This is a finishing touch on a topic that started in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>this post</strong></a>. In the original post we had a single neural network model scoring 0.790 on public LB. Here we have 12 different models (a mix of NN, LGB and logistic regression; scores are 0.78-0.795 on public LB) that are stacked together by another neural network. The model scores 0.799 on public LB, and t-SNE of its activations shows much better separation between the two classes. Notice that both tips of the worm have increased density of points, as a stack of models predicts many samples with extreme certainty.</p>\n<p><img src=\"https://i.ibb.co/9rGwK67/t-SNE-keras-training-06.png\" alt=\"t-SNE plot\"></p>\n<p><strong>EDIT</strong>: There is a plot <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271\" target=\"_blank\"><strong>here</strong></a> showing how PCA interprets this dataset.</p>",
      "rawMarkdown": "This is a finishing touch on a topic that started in [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278). In the original post we had a single neural network model scoring 0.790 on public LB. Here we have 12 different models (a mix of NN, LGB and logistic regression; scores are 0.78-0.795 on public LB) that are stacked together by another neural network. The model scores 0.799 on public LB, and t-SNE of its activations shows much better separation between the two classes. Notice that both tips of the worm have increased density of points, as a stack of models predicts many samples with extreme certainty.\n\n![t-SNE plot](https://i.ibb.co/9rGwK67/t-SNE-keras-training-06.png)\n\n**EDIT**: There is a plot [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271 ) showing how PCA interprets this dataset.",
      "votes": 50
    },
    {
      "id": 1874671,
      "postDate": "2022-07-28T12:22:16.187Z",
      "content": "<p>I'm new to data science and understand almost nothing of what you're saying, but these visuals are lovely, in all their variations. <br>\nThank you for the inspiration.</p>",
      "rawMarkdown": "I'm new to data science and understand almost nothing of what you're saying, but these visuals are lovely, in all their variations. \nThank you for the inspiration.",
      "votes": 3,
      "replies": [
        {
          "id": 1874934,
          "postDate": "2022-07-28T15:53:41.310Z",
          "content": "<p>There are 3 related posts I wrote on this subject, and they are cross-linked. It may help if you read them all, and feel free to ask if something specific is still unclear.</p>",
          "rawMarkdown": "There are 3 related posts I wrote on this subject, and they are cross-linked. It may help if you read them all, and feel free to ask if something specific is still unclear."
        },
        {
          "id": 1875303,
          "postDate": "2022-07-28T22:01:18.033Z",
          "content": "<p>I'll do that. Thank you.</p>",
          "rawMarkdown": "I'll do that. Thank you."
        }
      ]
    },
    {
      "id": 1873136,
      "postDate": "2022-07-27T13:07:06.230Z",
      "content": "<p>I did the same exercise and got few strange paintings :) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fa78dafbd0d96e43403ca7ee6ba25468e%2Ftsne_fig1.png?generation=1658926473095519&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2F838fe24580cd0bdff243481e10c85a2a%2Fumap_fig1.png?generation=1658926566640056&amp;alt=media\" alt=\"\"></p>\n<p>Not sure why the t-SNE look so different from Tillli's plot and how to interpret this, but is really fun to play with it. </p>\n<p>PS: I used just one fold (~91k customers) and default params for t-SNE, UMAP<br>\nPS: CV 0.798 / LB 0.797+</p>",
      "rawMarkdown": "I did the same exercise and got few strange paintings :) \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fa78dafbd0d96e43403ca7ee6ba25468e%2Ftsne_fig1.png?generation=1658926473095519&alt=media)\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2F838fe24580cd0bdff243481e10c85a2a%2Fumap_fig1.png?generation=1658926566640056&alt=media)\n\n\n\nNot sure why the t-SNE look so different from Tillli's plot and how to interpret this, but is really fun to play with it. \n\nPS: I used just one fold (~91k customers) and default params for t-SNE, UMAP\nPS: CV 0.798 / LB 0.797+\n",
      "votes": 4,
      "replies": [
        {
          "id": 1873588,
          "postDate": "2022-07-27T17:31:32.740Z",
          "content": "<p>t-SNE is very sensitive to it's hyperparameters. Here is a great article I found some time ago:<br>\n<a href=\"https://distill.pub/2016/misread-tsne/\" target=\"_blank\">https://distill.pub/2016/misread-tsne/</a></p>",
          "rawMarkdown": "t-SNE is very sensitive to it's hyperparameters. Here is a great article I found some time ago:\nhttps://distill.pub/2016/misread-tsne/",
          "votes": 3
        },
        {
          "id": 1873697,
          "postDate": "2022-07-27T20:25:31.320Z",
          "content": "<p>What <a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> said is true, except that I wouldn't use a plural. In my experience, the only parameter that truly matters is perplexity. Smaller values of perplexity tend to make plots where the dots are diffuse and form squiggles, kind of like in your plot. My original plot (not from this page, but <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>here</strong></a>) was done with <code>perplexity=20</code>. If I use the same dataset but set <code>perplexity=5</code>, this is what comes out as a result.</p>\n<p><img src=\"https://i.ibb.co/ZJDQZtN/t-SNE-keras-training-02a.png\" alt=\"t-SNE plot\"></p>\n<p>So I suggest you try the same t-SNE implementation I referenced in the original post, and play with perplexity until you get the embedding that makes most sense. Or just leave it as is, because this is still very nice graphics 😊</p>",
          "rawMarkdown": "What @thomasmeiner said is true, except that I wouldn't use a plural. In my experience, the only parameter that truly matters is perplexity. Smaller values of perplexity tend to make plots where the dots are diffuse and form squiggles, kind of like in your plot. My original plot (not from this page, but [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278)) was done with `perplexity=20`. If I use the same dataset but set `perplexity=5`, this is what comes out as a result.\n\n![t-SNE plot](https://i.ibb.co/ZJDQZtN/t-SNE-keras-training-02a.png)\n\nSo I suggest you try the same t-SNE implementation I referenced in the original post, and play with perplexity until you get the embedding that makes most sense. Or just leave it as is, because this is still very nice graphics 😊",
          "votes": 2
        },
        {
          "id": 1873707,
          "postDate": "2022-07-27T20:33:57.987Z",
          "content": "<p>yes exactly, perplexity is the key param. I played with some values and graphs change a lot. <br>\nFor reference my plot above is with RAPIDS cuml t-SNE(perplexity=30) which is the default value from what I see.</p>",
          "rawMarkdown": "yes exactly, perplexity is the key param. I played with some values and graphs change a lot. \nFor reference my plot above is with RAPIDS cuml t-SNE(perplexity=30) which is the default value from what I see.",
          "votes": 1
        },
        {
          "id": 1873717,
          "postDate": "2022-07-27T20:52:04.900Z",
          "content": "<p>Don't know about RAPIDS implementation, but my experience is that sklearn's t-SNE is inferior both in terms of speed and results. Again, I recommend <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>openTSNE</strong></a>.</p>",
          "rawMarkdown": "Don't know about RAPIDS implementation, but my experience is that sklearn's t-SNE is inferior both in terms of speed and results. Again, I recommend [**openTSNE**](https://github.com/pavlin-policar/openTSNE).",
          "votes": 1
        },
        {
          "id": 1873723,
          "postDate": "2022-07-27T21:04:11.190Z",
          "content": "<p>Yes it is perplexity. The article linked above shows that as well.</p>",
          "rawMarkdown": "Yes it is perplexity. The article linked above shows that as well.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1871340,
      "postDate": "2022-07-26T08:11:11.627Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>\n<p>Very nice! <br>\nBTW, I have found that using different colour maps can also create <a href=\"https://www.kaggle.com/code/carlmcbrideellis/some-pretty-t-sne-plots\" target=\"_blank\">some pretty t-SNE plots</a>…</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @tilii7 \n\nVery nice! \nBTW, I have found that using different colour maps can also create [some pretty t-SNE plots](https://www.kaggle.com/code/carlmcbrideellis/some-pretty-t-sne-plots)...\n\nAll the best,\ncarl",
      "votes": 4,
      "replies": [
        {
          "id": 1871426,
          "postDate": "2022-07-26T09:07:00.377Z",
          "content": "<p>Since you want more color, here is the same image as above but colored according to network predictions rather than actual class values. Even though t-SNE never sees the actual class labels or neural network predictions, its embedding matches the distribution of predictions.</p>\n<p><img src=\"https://i.ibb.co/F4Wnh1k/t-SNE-keras-training-08a.png\" alt=\"t-SNE plot\"></p>",
          "rawMarkdown": "Since you want more color, here is the same image as above but colored according to network predictions rather than actual class values. Even though t-SNE never sees the actual class labels or neural network predictions, its embedding matches the distribution of predictions.\n\n![t-SNE plot](https://i.ibb.co/F4Wnh1k/t-SNE-keras-training-08a.png)",
          "votes": 5
        }
      ]
    },
    {
      "id": 1897397,
      "postDate": "2022-08-13T17:41:02.760Z",
      "content": "<p>Wow,  is very very cool .</p>",
      "rawMarkdown": "Wow,  is very very cool .",
      "votes": 1
    },
    {
      "id": 1883764,
      "postDate": "2022-08-04T04:57:59.173Z",
      "content": "<p>Wow, this is beautiful</p>",
      "rawMarkdown": "Wow, this is beautiful",
      "votes": 1
    },
    {
      "id": 1882853,
      "postDate": "2022-08-03T13:01:29.997Z",
      "content": "<p>This is art! I can't think of any other words to compliment you.👍</p>",
      "rawMarkdown": "This is art! I can't think of any other words to compliment you.👍",
      "votes": 1
    },
    {
      "id": 1882197,
      "postDate": "2022-08-03T05:57:23.510Z",
      "content": "<p>Great plot! I have to improve my skills in neural network</p>",
      "rawMarkdown": "Great plot! I have to improve my skills in neural network",
      "votes": 1
    },
    {
      "id": 1882071,
      "postDate": "2022-08-03T04:07:50.680Z",
      "content": "<p>wow, great job</p>",
      "rawMarkdown": "wow, great job",
      "votes": 1
    },
    {
      "id": 1875807,
      "postDate": "2022-07-29T10:01:22.173Z",
      "content": "<p>data is art indeed! +2</p>",
      "rawMarkdown": "data is art indeed! +2",
      "votes": 1
    },
    {
      "id": 1873819,
      "postDate": "2022-07-27T23:20:02.877Z",
      "content": "<p>Nice one 😂!<br>\nIt is really interesting.. why this happen?</p>",
      "rawMarkdown": "Nice one 😂!\nIt is really interesting.. why this happen?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1874007,
          "postDate": "2022-07-28T02:35:49.913Z",
          "content": "<p>I am not sure what exactly you are asking.</p>\n<p>If you are asking why data classes separate, it is because the underlying neural network has already learned the pattern that makes them different. All that is left is for this vector of 50 numbers to be fed into a sigmoid layer, which will convert the information into sigmoidally-distributed probabilities. Instead, we feed the information into t-SNE, which will reduce the dimensionality from 50 to 2, but still maintain nice data separation.</p>\n<p>If you are asking why the overall plot looks like a worm, it is kind of explained in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271\" target=\"_blank\"><strong>this comment</strong></a>. Briefly, t-SNE tries to keep locally similar samples together, without necessarily maintaining global distance between data points. Since the perplexity parameter was set at 20, in plain terms it has to make sure that each point has at least another 20 points nearby. That allows for quite a big stretch of data points as they don't have to be densely packed, and eventually the whole thing winds around like a worm. Changing the perplexity parameter to a large value, say 1000, will now require that each point has another 1000 points nearby. That will make the plot more compact because data points can't be spread so much, and that will likely remove the \"worminess.\" See below the same data as above in the main plot, but with <code>perplexity=1000</code>. Note that the scale in this plot is <code>[-60, 60]</code> while it is <code>[-300, 300]</code> above, which means that data spread is much more compact.</p>\n<p><img src=\"https://i.ibb.co/8mqQGPD/t-SNE-keras-training-09.png\" alt=\"t-SNE plot\"></p>",
          "rawMarkdown": "I am not sure what exactly you are asking.\n\nIf you are asking why data classes separate, it is because the underlying neural network has already learned the pattern that makes them different. All that is left is for this vector of 50 numbers to be fed into a sigmoid layer, which will convert the information into sigmoidally-distributed probabilities. Instead, we feed the information into t-SNE, which will reduce the dimensionality from 50 to 2, but still maintain nice data separation.\n\nIf you are asking why the overall plot looks like a worm, it is kind of explained in [**this comment**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271). Briefly, t-SNE tries to keep locally similar samples together, without necessarily maintaining global distance between data points. Since the perplexity parameter was set at 20, in plain terms it has to make sure that each point has at least another 20 points nearby. That allows for quite a big stretch of data points as they don't have to be densely packed, and eventually the whole thing winds around like a worm. Changing the perplexity parameter to a large value, say 1000, will now require that each point has another 1000 points nearby. That will make the plot more compact because data points can't be spread so much, and that will likely remove the \"worminess.\" See below the same data as above in the main plot, but with `perplexity=1000`. Note that the scale in this plot is `[-60, 60]` while it is `[-300, 300]` above, which means that data spread is much more compact.\n\n![t-SNE plot](https://i.ibb.co/8mqQGPD/t-SNE-keras-training-09.png)\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 1873181,
      "postDate": "2022-07-27T13:40:19.650Z",
      "content": "<p>Incredibly interesting, thank you! </p>",
      "rawMarkdown": "Incredibly interesting, thank you! ~~Have you seen the worm on a string memes by chance?~~",
      "votes": 1
    },
    {
      "id": 1896845,
      "postDate": "2022-08-13T08:02:36.237Z",
      "content": "<p>wow, Great job</p>",
      "rawMarkdown": "wow, Great job"
    },
    {
      "id": 1875838,
      "postDate": "2022-07-29T10:37:02.380Z",
      "content": "<p>Wow,  the picture is very very cool .</p>\n<p>Data is Art,  and BTW,  my startup is named BoolArt,  lol .   </p>",
      "rawMarkdown": "Wow,  the picture is very very cool .\n\nData is Art,  and BTW,  my startup is named BoolArt,  lol .   ",
      "votes": 2
    },
    {
      "id": 1874502,
      "postDate": "2022-07-28T10:12:14.400Z",
      "content": "<p>Data is creative !</p>",
      "rawMarkdown": "Data is creative !",
      "votes": 2
    },
    {
      "id": 1872460,
      "postDate": "2022-07-27T02:01:12.110Z",
      "content": "<p>beautiful plot <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
      "rawMarkdown": "beautiful plot @tilii7 ",
      "votes": 2
    },
    {
      "id": 1872299,
      "postDate": "2022-07-26T19:41:42.447Z",
      "content": "<p>very artistic data</p>",
      "rawMarkdown": "very artistic data",
      "votes": 2
    },
    {
      "id": 1871659,
      "postDate": "2022-07-26T11:56:44.333Z",
      "content": "<p>data is art indeed! +1</p>",
      "rawMarkdown": "data is art indeed! +1",
      "votes": 2
    },
    {
      "id": 1884086,
      "postDate": "2022-08-04T09:24:30.097Z",
      "content": "<p>wow, great job</p>",
      "rawMarkdown": "wow, great job",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1883683,
      "postDate": "2022-08-04T02:47:29.267Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1873431,
      "postDate": "2022-07-27T15:52:38.490Z",
      "content": "<p>thanks sir <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
      "rawMarkdown": "thanks sir @tilii7 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1874671,
      "author_name": "SKLasky",
      "author_url": "",
      "post_date": "2022-07-28T12:22:16.187000",
      "content": "<p>I'm new to data science and understand almost nothing of what you're saying, but these visuals are lovely, in all their variations. <br>\nThank you for the inspiration.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1874934,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-07-28T15:53:41.310000",
          "content": "<p>There are 3 related posts I wrote on this subject, and they are cross-linked. It may help if you read them all, and feel free to ask if something specific is still unclear.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1875303,
          "author_name": "SKLasky",
          "author_url": "",
          "post_date": "2022-07-28T22:01:18.033000",
          "content": "<p>I'll do that. Thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1873136,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2022-07-27T13:07:06.230000",
      "content": "<p>I did the same exercise and got few strange paintings :) </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fa78dafbd0d96e43403ca7ee6ba25468e%2Ftsne_fig1.png?generation=1658926473095519&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2F838fe24580cd0bdff243481e10c85a2a%2Fumap_fig1.png?generation=1658926566640056&amp;alt=media\" alt=\"\"></p>\n<p>Not sure why the t-SNE look so different from Tillli's plot and how to interpret this, but is really fun to play with it. </p>\n<p>PS: I used just one fold (~91k customers) and default params for t-SNE, UMAP<br>\nPS: CV 0.798 / LB 0.797+</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1873588,
          "author_name": "Thomas Meißner",
          "author_url": "",
          "post_date": "2022-07-27T17:31:32.740000",
          "content": "<p>t-SNE is very sensitive to it's hyperparameters. Here is a great article I found some time ago:<br>\n<a href=\"https://distill.pub/2016/misread-tsne/\" target=\"_blank\">https://distill.pub/2016/misread-tsne/</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1873697,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-07-27T20:25:31.320000",
          "content": "<p>What <a href=\"https://www.kaggle.com/thomasmeiner\" target=\"_blank\">@thomasmeiner</a> said is true, except that I wouldn't use a plural. In my experience, the only parameter that truly matters is perplexity. Smaller values of perplexity tend to make plots where the dots are diffuse and form squiggles, kind of like in your plot. My original plot (not from this page, but <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278\" target=\"_blank\"><strong>here</strong></a>) was done with <code>perplexity=20</code>. If I use the same dataset but set <code>perplexity=5</code>, this is what comes out as a result.</p>\n<p><img src=\"https://i.ibb.co/ZJDQZtN/t-SNE-keras-training-02a.png\" alt=\"t-SNE plot\"></p>\n<p>So I suggest you try the same t-SNE implementation I referenced in the original post, and play with perplexity until you get the embedding that makes most sense. Or just leave it as is, because this is still very nice graphics 😊</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1873707,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2022-07-27T20:33:57.987000",
          "content": "<p>yes exactly, perplexity is the key param. I played with some values and graphs change a lot. <br>\nFor reference my plot above is with RAPIDS cuml t-SNE(perplexity=30) which is the default value from what I see.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1873717,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-07-27T20:52:04.900000",
          "content": "<p>Don't know about RAPIDS implementation, but my experience is that sklearn's t-SNE is inferior both in terms of speed and results. Again, I recommend <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>openTSNE</strong></a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1873723,
          "author_name": "Thomas Meißner",
          "author_url": "",
          "post_date": "2022-07-27T21:04:11.190000",
          "content": "<p>Yes it is perplexity. The article linked above shows that as well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1871340,
      "author_name": "Carl McBride Ellis",
      "author_url": "",
      "post_date": "2022-07-26T08:11:11.627000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>\n<p>Very nice! <br>\nBTW, I have found that using different colour maps can also create <a href=\"https://www.kaggle.com/code/carlmcbrideellis/some-pretty-t-sne-plots\" target=\"_blank\">some pretty t-SNE plots</a>…</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1871426,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-07-26T09:07:00.377000",
          "content": "<p>Since you want more color, here is the same image as above but colored according to network predictions rather than actual class values. Even though t-SNE never sees the actual class labels or neural network predictions, its embedding matches the distribution of predictions.</p>\n<p><img src=\"https://i.ibb.co/F4Wnh1k/t-SNE-keras-training-08a.png\" alt=\"t-SNE plot\"></p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1897397,
      "author_name": "Taylorll",
      "author_url": "",
      "post_date": "2022-08-13T17:41:02.760000",
      "content": "<p>Wow,  is very very cool .</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1883764,
      "author_name": "Will",
      "author_url": "",
      "post_date": "2022-08-04T04:57:59.173000",
      "content": "<p>Wow, this is beautiful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1882853,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-03T13:01:29.997000",
      "content": "<p>This is art! I can't think of any other words to compliment you.👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1882197,
      "author_name": "Agustin222",
      "author_url": "",
      "post_date": "2022-08-03T05:57:23.510000",
      "content": "<p>Great plot! I have to improve my skills in neural network</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1882071,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-03T04:07:50.680000",
      "content": "<p>wow, great job</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1875807,
      "author_name": "Kuzma Peng",
      "author_url": "",
      "post_date": "2022-07-29T10:01:22.173000",
      "content": "<p>data is art indeed! +2</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1873819,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-27T23:20:02.877000",
      "content": "<p>Nice one 😂!<br>\nIt is really interesting.. why this happen?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1874007,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-07-28T02:35:49.913000",
          "content": "<p>I am not sure what exactly you are asking.</p>\n<p>If you are asking why data classes separate, it is because the underlying neural network has already learned the pattern that makes them different. All that is left is for this vector of 50 numbers to be fed into a sigmoid layer, which will convert the information into sigmoidally-distributed probabilities. Instead, we feed the information into t-SNE, which will reduce the dimensionality from 50 to 2, but still maintain nice data separation.</p>\n<p>If you are asking why the overall plot looks like a worm, it is kind of explained in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271\" target=\"_blank\"><strong>this comment</strong></a>. Briefly, t-SNE tries to keep locally similar samples together, without necessarily maintaining global distance between data points. Since the perplexity parameter was set at 20, in plain terms it has to make sure that each point has at least another 20 points nearby. That allows for quite a big stretch of data points as they don't have to be densely packed, and eventually the whole thing winds around like a worm. Changing the perplexity parameter to a large value, say 1000, will now require that each point has another 1000 points nearby. That will make the plot more compact because data points can't be spread so much, and that will likely remove the \"worminess.\" See below the same data as above in the main plot, but with <code>perplexity=1000</code>. Note that the scale in this plot is <code>[-60, 60]</code> while it is <code>[-300, 300]</code> above, which means that data spread is much more compact.</p>\n<p><img src=\"https://i.ibb.co/8mqQGPD/t-SNE-keras-training-09.png\" alt=\"t-SNE plot\"></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1873181,
      "author_name": "Bella Pisani",
      "author_url": "",
      "post_date": "2022-07-27T13:40:19.650000",
      "content": "<p>Incredibly interesting, thank you! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1896845,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T08:02:36.237000",
      "content": "<p>wow, Great job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1875838,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-07-29T10:37:02.380000",
      "content": "<p>Wow,  the picture is very very cool .</p>\n<p>Data is Art,  and BTW,  my startup is named BoolArt,  lol .   </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1874502,
      "author_name": "Shrijayan",
      "author_url": "",
      "post_date": "2022-07-28T10:12:14.400000",
      "content": "<p>Data is creative !</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1872460,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-07-27T02:01:12.110000",
      "content": "<p>beautiful plot <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1872299,
      "author_name": "Orkhan Suleymanli",
      "author_url": "",
      "post_date": "2022-07-26T19:41:42.447000",
      "content": "<p>very artistic data</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1871659,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2022-07-26T11:56:44.333000",
      "content": "<p>data is art indeed! +1</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1884086,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-04T09:24:30.097000",
      "content": "<p>wow, great job</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1883683,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-04T02:47:29.267000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1873431,
      "author_name": "Bilal Suppal",
      "author_url": "",
      "post_date": "2022-07-27T15:52:38.490000",
      "content": "<p>thanks sir <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1871311": "This is a finishing touch on a topic that started in [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278). In the original post we had a single neural network model scoring 0.790 on public LB. Here we have 12 different models (a mix of NN, LGB and logistic regression; scores are 0.78-0.795 on public LB) that are stacked together by another neural network. The model scores 0.799 on public LB, and t-SNE of its activations shows much better separation between the two classes. Notice that both tips of the worm have increased density of points, as a stack of models predicts many samples with extreme certainty.\n\n![t-SNE plot](https://i.ibb.co/9rGwK67/t-SNE-keras-training-06.png)\n\n**EDIT**: There is a plot [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339278#1872271 ) showing how PCA interprets this dataset.",
    "1874671": "I'm new to data science and understand almost nothing of what you're saying, but these visuals are lovely, in all their variations. \nThank you for the inspiration.",
    "1873136": "I did the same exercise and got few strange paintings :) \n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2Fa78dafbd0d96e43403ca7ee6ba25468e%2Ftsne_fig1.png?generation=1658926473095519&alt=media)\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2596066%2F838fe24580cd0bdff243481e10c85a2a%2Fumap_fig1.png?generation=1658926566640056&alt=media)\n\n\n\nNot sure why the t-SNE look so different from Tillli's plot and how to interpret this, but is really fun to play with it. \n\nPS: I used just one fold (~91k customers) and default params for t-SNE, UMAP\nPS: CV 0.798 / LB 0.797+\n",
    "1871340": "Dear @tilii7 \n\nVery nice! \nBTW, I have found that using different colour maps can also create [some pretty t-SNE plots](https://www.kaggle.com/code/carlmcbrideellis/some-pretty-t-sne-plots)...\n\nAll the best,\ncarl",
    "1897397": "Wow,  is very very cool .",
    "1883764": "Wow, this is beautiful",
    "1882853": "This is art! I can't think of any other words to compliment you.👍",
    "1882197": "Great plot! I have to improve my skills in neural network",
    "1882071": "wow, great job",
    "1875807": "data is art indeed! +2",
    "1873819": "Nice one 😂!\nIt is really interesting.. why this happen?\n",
    "1873181": "Incredibly interesting, thank you! ~~Have you seen the worm on a string memes by chance?~~",
    "1896845": "wow, Great job",
    "1875838": "Wow,  the picture is very very cool .\n\nData is Art,  and BTW,  my startup is named BoolArt,  lol .   ",
    "1874502": "Data is creative !",
    "1872460": "beautiful plot @tilii7 ",
    "1872299": "very artistic data",
    "1871659": "data is art indeed! +1",
    "1884086": "wow, great job",
    "1883683": "",
    "1873431": "thanks sir @tilii7 "
  }
}