{
  "id": 342679,
  "title": "t-SNE learns to separate data classess -- a short movie",
  "url": "/competitions/amex-default-prediction/discussion/342679",
  "author_name": "Tilii",
  "post_date": "2022-08-08T10:33:20.291000",
  "votes": 28,
  "comment_count": 19,
  "views": 0,
  "content": "<p>This is similar to what I already discussed <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726\" target=\"_blank\"><strong>here</strong></a>. This time we have a little bit better stacking neural network, and a movie showing the whole process of how data embedding happens with t-SNE. I think it is a good approximation for how the neural network makes the final classification step. Shown at the bottom is the movie's final frame if you want to inspect it at higher resolution.</p>\n<p><img src=\"https://i.ibb.co/MMbDJqV/plot-t-SNE-anim-01.gif\" alt=\"t-SNE learning\"></p>\n<p><img src=\"https://i.ibb.co/Vm9qgvG/t-SNE-keras-training-anim-01.png\" alt=\"t-SNE plot\"></p>",
  "messages": [
    {
      "id": 1889671,
      "postDate": "2022-08-08T10:33:20.290Z",
      "content": "<p>This is similar to what I already discussed <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726\" target=\"_blank\"><strong>here</strong></a>. This time we have a little bit better stacking neural network, and a movie showing the whole process of how data embedding happens with t-SNE. I think it is a good approximation for how the neural network makes the final classification step. Shown at the bottom is the movie's final frame if you want to inspect it at higher resolution.</p>\n<p><img src=\"https://i.ibb.co/MMbDJqV/plot-t-SNE-anim-01.gif\" alt=\"t-SNE learning\"></p>\n<p><img src=\"https://i.ibb.co/Vm9qgvG/t-SNE-keras-training-anim-01.png\" alt=\"t-SNE plot\"></p>",
      "rawMarkdown": "This is similar to what I already discussed [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726). This time we have a little bit better stacking neural network, and a movie showing the whole process of how data embedding happens with t-SNE. I think it is a good approximation for how the neural network makes the final classification step. Shown at the bottom is the movie's final frame if you want to inspect it at higher resolution.\n\n![t-SNE learning](https://i.ibb.co/MMbDJqV/plot-t-SNE-anim-01.gif)\n\n![t-SNE plot](https://i.ibb.co/Vm9qgvG/t-SNE-keras-training-anim-01.png)\n",
      "votes": 28
    },
    {
      "id": 1903290,
      "postDate": "2022-08-17T09:22:39.040Z",
      "content": "<p>WOW, So beautiful</p>",
      "rawMarkdown": "WOW, So beautiful",
      "votes": 1
    },
    {
      "id": 1898350,
      "postDate": "2022-08-14T14:29:32.193Z",
      "content": "<p>That is so cool!</p>",
      "rawMarkdown": "That is so cool!",
      "votes": 1
    },
    {
      "id": 1896862,
      "postDate": "2022-08-13T08:19:29.353Z",
      "content": "<p>that's really cool!</p>",
      "rawMarkdown": "that's really cool!",
      "votes": 1
    },
    {
      "id": 1895468,
      "postDate": "2022-08-12T06:48:43.100Z",
      "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Fantastic visual &amp; frames of manifold learning.. thanks for sharing!</p>",
      "rawMarkdown": "@tilii7 Fantastic visual & frames of manifold learning.. thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1893457,
      "postDate": "2022-08-10T20:37:46.253Z",
      "content": "<p>That visualization is incredible, do you mind sharing a sample code, it would be interesting to see how data is being classified in other cases. Does the visualization take long to run on say a sample set of 100,000 rows with 15 features?  </p>",
      "rawMarkdown": "That visualization is incredible, do you mind sharing a sample code, it would be interesting to see how data is being classified in other cases. Does the visualization take long to run on say a sample set of 100,000 rows with 15 features?  ",
      "votes": 1,
      "replies": [
        {
          "id": 1893523,
          "postDate": "2022-08-10T22:14:50.777Z",
          "content": "<p>First thing first: this is only an approximation of how data is classified. The input data consists of penultimate layer activations (50 features), which in case of the neural network go into the final sigmoid layer where the classification is done. What t-SNE does with these activations is not the same: it will reduce data dimensionality from 50 to 2, and spread the points around the plot in such a way that immediate sample neighbors are close to each other, but their global positioning may not necessarily be representative of their true distances. It just so happens that during the t-SNE process some interesting trajectories are created, and the final result is often illustrative as well.</p>\n<p>It doesn't take long. <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>openTSNE</strong></a> is multithreaded, so for this sample (~91K points x 50 features) and with 10 CPUs it takes about half an hour to do the embedding and create the animation. The input data is obtained by cutting out the last neural network layer (the sigmoid layer) and doing predictions, which generates a 50-vector prediction for each data point since the penultimate layer has 50 neurons. You can choose a smaller number of neurons in that layer, so the embedding could be even faster. Once the input data is available, openTSNE takes over and it is pretty straightforward. I use perplexity of 40-100 (100 for the above animation) and there is an explanation on their website how to create an animation of the whole embedding. I typically cut out the early exaggeration phase (also explained) because not much happens visually during that time.</p>",
          "rawMarkdown": "First thing first: this is only an approximation of how data is classified. The input data consists of penultimate layer activations (50 features), which in case of the neural network go into the final sigmoid layer where the classification is done. What t-SNE does with these activations is not the same: it will reduce data dimensionality from 50 to 2, and spread the points around the plot in such a way that immediate sample neighbors are close to each other, but their global positioning may not necessarily be representative of their true distances. It just so happens that during the t-SNE process some interesting trajectories are created, and the final result is often illustrative as well.\n\nIt doesn't take long. [**openTSNE**](https://github.com/pavlin-policar/openTSNE) is multithreaded, so for this sample (~91K points x 50 features) and with 10 CPUs it takes about half an hour to do the embedding and create the animation. The input data is obtained by cutting out the last neural network layer (the sigmoid layer) and doing predictions, which generates a 50-vector prediction for each data point since the penultimate layer has 50 neurons. You can choose a smaller number of neurons in that layer, so the embedding could be even faster. Once the input data is available, openTSNE takes over and it is pretty straightforward. I use perplexity of 40-100 (100 for the above animation) and there is an explanation on their website how to create an animation of the whole embedding. I typically cut out the early exaggeration phase (also explained) because not much happens visually during that time.",
          "votes": 3
        },
        {
          "id": 1899656,
          "postDate": "2022-08-15T12:31:32.207Z",
          "content": "<p>Do you have an example how to use open tsne ?</p>",
          "rawMarkdown": "Do you have an example how to use open tsne ?"
        },
        {
          "id": 1899869,
          "postDate": "2022-08-15T14:58:24.310Z",
          "content": "<p>I don't have any code I could share, because most of my t-SNE code is written for other purposes. It is very simple to get it going, and they have extensive documentation. Any kind of tabular data works (class/target values should be excluded), and it is a good idea to normalize the values. For &lt;50,000 data points, what I show below should work fine. For &gt;50,000 data points, I suggest changing <code>early_exaggeration_iter=2000</code> and <code>n_iter=5000</code>. You may want to experiment with perplexity values for any specific application, though for DNA embedding from 4n frequencies 20-40 works fine. Also, openTSNE will happily use all the CPUs you have, so you may want to increase <code>n_jobs</code> depending on your system.</p>\n<pre><code>from openTSNE import TSNE\n\n# load tabular data into train_data, normalize if needed\n\ntsne = TSNE(\n    perplexity=40,\n    metric='cosine',\n    n_jobs=10,\n    learning_rate='auto',\n    negative_gradient_method='bh',\n    theta=0.5,\n    initialization='spectral',\n    early_exaggeration_iter=1000,\n    early_exaggeration=18.0,\n    n_iter=3000,\n    initial_momentum=0.5,\n    final_momentum=0.8,\n    verbose=1)\n\nembedding = tsne.fit(train_data)\n</code></pre>",
          "rawMarkdown": "I don't have any code I could share, because most of my t-SNE code is written for other purposes. It is very simple to get it going, and they have extensive documentation. Any kind of tabular data works (class/target values should be excluded), and it is a good idea to normalize the values. For <50,000 data points, what I show below should work fine. For >50,000 data points, I suggest changing `early_exaggeration_iter=2000` and `n_iter=5000`. You may want to experiment with perplexity values for any specific application, though for DNA embedding from 4n frequencies 20-40 works fine. Also, openTSNE will happily use all the CPUs you have, so you may want to increase `n_jobs` depending on your system.\n\n```\nfrom openTSNE import TSNE\n    \n# load tabular data into train_data, normalize if needed\n    \ntsne = TSNE(\n    perplexity=40,\n    metric='cosine',\n    n_jobs=10,\n    learning_rate='auto',\n    negative_gradient_method='bh',\n    theta=0.5,\n    initialization='spectral',\n    early_exaggeration_iter=1000,\n    early_exaggeration=18.0,\n    n_iter=3000,\n    initial_momentum=0.5,\n    final_momentum=0.8,\n    verbose=1)\n    \nembedding = tsne.fit(train_data)\n```\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1892175,
      "postDate": "2022-08-09T22:56:57.660Z",
      "content": "<p>This is really interesting, check my new article on Devgenius classifying bacteria DNA segments using tsne : <br>\n<a href=\"https://medium.com/dev-genius/data-science-exploratory-analysis-code-blocks-used-by-data-engineers-ce4529b70672\" target=\"_blank\">https://medium.com/dev-genius/data-science-exploratory-analysis-code-blocks-used-by-data-engineers-ce4529b70672</a></p>",
      "rawMarkdown": "This is really interesting, check my new article on Devgenius classifying bacteria DNA segments using tsne : \nhttps://medium.com/dev-genius/data-science-exploratory-analysis-code-blocks-used-by-data-engineers-ce4529b70672",
      "votes": 1,
      "replies": [
        {
          "id": 1892317,
          "postDate": "2022-08-10T03:21:50.303Z",
          "content": "<p>Nicely done. You should know that sklearn's t-SNE implementation is really substandard for these kinds of jobs, and I think in general.</p>\n<p>Hundreds of microbial genomes can be separated from a complex mixture using <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>openTSNE</strong></a>. The highest cluster number in the image below is 351, and that's before the refinment.</p>\n<p><img src=\"https://i.ibb.co/QXQ1jvC/contigs-hdbscan-20k-5k-2-D-opent-SNE-clusters.png\" alt=\"t-SNE binning\"></p>",
          "rawMarkdown": "Nicely done. You should know that sklearn's t-SNE implementation is really substandard for these kinds of jobs, and I think in general.\n\nHundreds of microbial genomes can be separated from a complex mixture using [**openTSNE**](https://github.com/pavlin-policar/openTSNE). The highest cluster number in the image below is 351, and that's before the refinment.\n\n![t-SNE binning](https://i.ibb.co/QXQ1jvC/contigs-hdbscan-20k-5k-2-D-opent-SNE-clusters.png)",
          "votes": 2
        },
        {
          "id": 1893531,
          "postDate": "2022-08-10T22:24:42.850Z",
          "content": "<p>Do you know how can I extract the segments dataset after analysis ?</p>",
          "rawMarkdown": "Do you know how can I extract the segments dataset after analysis ?"
        },
        {
          "id": 1893592,
          "postDate": "2022-08-10T23:27:46.257Z",
          "content": "<p>I use scripts from the <a href=\"https://github.com/BinPro/CONCOCT\" target=\"_blank\"><strong>CONCOCT</strong></a> package, but that assumes cutting up DNA into pieces the way they do it. Relevant scripts are <code>extract_fasta_bins.py</code> and <code>merge_cutup_clustering.py</code>. You may want to go end-to-end with CONCOCT, though I get better results for complex datasets with openTSNE.</p>",
          "rawMarkdown": "I use scripts from the [**CONCOCT**](https://github.com/BinPro/CONCOCT) package, but that assumes cutting up DNA into pieces the way they do it. Relevant scripts are `extract_fasta_bins.py` and `merge_cutup_clustering.py`. You may want to go end-to-end with CONCOCT, though I get better results for complex datasets with openTSNE.",
          "votes": 1
        },
        {
          "id": 1893595,
          "postDate": "2022-08-10T23:34:07.947Z",
          "content": "<p>Indeed very interesting I will advise you to write an article on this 👍👍</p>",
          "rawMarkdown": "Indeed very interesting I will advise you to write an article on this 👍👍",
          "votes": 1
        },
        {
          "id": 1911904,
          "postDate": "2022-08-24T11:35:57.090Z",
          "content": "<p>Do you have a code example using concoct ? thanks for your help it is just for learning new methods 😃</p>",
          "rawMarkdown": "Do you have a code example using concoct ? thanks for your help it is just for learning new methods 😃"
        }
      ]
    },
    {
      "id": 1889801,
      "postDate": "2022-08-08T11:29:04.293Z",
      "content": "<p>so cool !!</p>",
      "rawMarkdown": "so cool !!",
      "votes": 1
    },
    {
      "id": 1892020,
      "postDate": "2022-08-09T19:50:24.703Z",
      "content": "<p>The animation is really cool! </p>",
      "rawMarkdown": "The animation is really cool! \n\n\n",
      "votes": 2
    },
    {
      "id": 1904167,
      "postDate": "2022-08-18T02:37:59.043Z",
      "content": "<p>This is too cool </p>",
      "rawMarkdown": "This is too cool "
    },
    {
      "id": 1892884,
      "postDate": "2022-08-10T11:59:25.670Z",
      "content": "<p>Thank you for the great answer I will try the openTsne, it seems a great way to find new antibiotics even for monkeypox.</p>",
      "rawMarkdown": "Thank you for the great answer I will try the openTsne, it seems a great way to find new antibiotics even for monkeypox."
    },
    {
      "id": 1900433,
      "postDate": "2022-08-16T03:18:50.413Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1903290,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-17T09:22:39.040000",
      "content": "<p>WOW, So beautiful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1898350,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T14:29:32.193000",
      "content": "<p>That is so cool!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1896862,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-13T08:19:29.353000",
      "content": "<p>that's really cool!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1895468,
      "author_name": "Ranja Sarkar (She/Her)",
      "author_url": "",
      "post_date": "2022-08-12T06:48:43.100000",
      "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> Fantastic visual &amp; frames of manifold learning.. thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1893457,
      "author_name": "JP",
      "author_url": "",
      "post_date": "2022-08-10T20:37:46.253000",
      "content": "<p>That visualization is incredible, do you mind sharing a sample code, it would be interesting to see how data is being classified in other cases. Does the visualization take long to run on say a sample set of 100,000 rows with 15 features?  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1893523,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-08-10T22:14:50.777000",
          "content": "<p>First thing first: this is only an approximation of how data is classified. The input data consists of penultimate layer activations (50 features), which in case of the neural network go into the final sigmoid layer where the classification is done. What t-SNE does with these activations is not the same: it will reduce data dimensionality from 50 to 2, and spread the points around the plot in such a way that immediate sample neighbors are close to each other, but their global positioning may not necessarily be representative of their true distances. It just so happens that during the t-SNE process some interesting trajectories are created, and the final result is often illustrative as well.</p>\n<p>It doesn't take long. <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>openTSNE</strong></a> is multithreaded, so for this sample (~91K points x 50 features) and with 10 CPUs it takes about half an hour to do the embedding and create the animation. The input data is obtained by cutting out the last neural network layer (the sigmoid layer) and doing predictions, which generates a 50-vector prediction for each data point since the penultimate layer has 50 neurons. You can choose a smaller number of neurons in that layer, so the embedding could be even faster. Once the input data is available, openTSNE takes over and it is pretty straightforward. I use perplexity of 40-100 (100 for the above animation) and there is an explanation on their website how to create an animation of the whole embedding. I typically cut out the early exaggeration phase (also explained) because not much happens visually during that time.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1899656,
          "author_name": "philBoaz",
          "author_url": "",
          "post_date": "2022-08-15T12:31:32.207000",
          "content": "<p>Do you have an example how to use open tsne ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1899869,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-08-15T14:58:24.310000",
          "content": "<p>I don't have any code I could share, because most of my t-SNE code is written for other purposes. It is very simple to get it going, and they have extensive documentation. Any kind of tabular data works (class/target values should be excluded), and it is a good idea to normalize the values. For &lt;50,000 data points, what I show below should work fine. For &gt;50,000 data points, I suggest changing <code>early_exaggeration_iter=2000</code> and <code>n_iter=5000</code>. You may want to experiment with perplexity values for any specific application, though for DNA embedding from 4n frequencies 20-40 works fine. Also, openTSNE will happily use all the CPUs you have, so you may want to increase <code>n_jobs</code> depending on your system.</p>\n<pre><code>from openTSNE import TSNE\n\n# load tabular data into train_data, normalize if needed\n\ntsne = TSNE(\n    perplexity=40,\n    metric='cosine',\n    n_jobs=10,\n    learning_rate='auto',\n    negative_gradient_method='bh',\n    theta=0.5,\n    initialization='spectral',\n    early_exaggeration_iter=1000,\n    early_exaggeration=18.0,\n    n_iter=3000,\n    initial_momentum=0.5,\n    final_momentum=0.8,\n    verbose=1)\n\nembedding = tsne.fit(train_data)\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1892175,
      "author_name": "philBoaz",
      "author_url": "",
      "post_date": "2022-08-09T22:56:57.660000",
      "content": "<p>This is really interesting, check my new article on Devgenius classifying bacteria DNA segments using tsne : <br>\n<a href=\"https://medium.com/dev-genius/data-science-exploratory-analysis-code-blocks-used-by-data-engineers-ce4529b70672\" target=\"_blank\">https://medium.com/dev-genius/data-science-exploratory-analysis-code-blocks-used-by-data-engineers-ce4529b70672</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1892317,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-08-10T03:21:50.303000",
          "content": "<p>Nicely done. You should know that sklearn's t-SNE implementation is really substandard for these kinds of jobs, and I think in general.</p>\n<p>Hundreds of microbial genomes can be separated from a complex mixture using <a href=\"https://github.com/pavlin-policar/openTSNE\" target=\"_blank\"><strong>openTSNE</strong></a>. The highest cluster number in the image below is 351, and that's before the refinment.</p>\n<p><img src=\"https://i.ibb.co/QXQ1jvC/contigs-hdbscan-20k-5k-2-D-opent-SNE-clusters.png\" alt=\"t-SNE binning\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1893531,
          "author_name": "philBoaz",
          "author_url": "",
          "post_date": "2022-08-10T22:24:42.850000",
          "content": "<p>Do you know how can I extract the segments dataset after analysis ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1893592,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2022-08-10T23:27:46.257000",
          "content": "<p>I use scripts from the <a href=\"https://github.com/BinPro/CONCOCT\" target=\"_blank\"><strong>CONCOCT</strong></a> package, but that assumes cutting up DNA into pieces the way they do it. Relevant scripts are <code>extract_fasta_bins.py</code> and <code>merge_cutup_clustering.py</code>. You may want to go end-to-end with CONCOCT, though I get better results for complex datasets with openTSNE.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1893595,
          "author_name": "philBoaz",
          "author_url": "",
          "post_date": "2022-08-10T23:34:07.947000",
          "content": "<p>Indeed very interesting I will advise you to write an article on this 👍👍</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1911904,
          "author_name": "philBoaz",
          "author_url": "",
          "post_date": "2022-08-24T11:35:57.090000",
          "content": "<p>Do you have a code example using concoct ? thanks for your help it is just for learning new methods 😃</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1889801,
      "author_name": "fengfeng",
      "author_url": "",
      "post_date": "2022-08-08T11:29:04.293000",
      "content": "<p>so cool !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1892020,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-08-09T19:50:24.703000",
      "content": "<p>The animation is really cool! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1904167,
      "author_name": "Pascal Yamlome",
      "author_url": "",
      "post_date": "2022-08-18T02:37:59.043000",
      "content": "<p>This is too cool </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1892884,
      "author_name": "philBoaz",
      "author_url": "",
      "post_date": "2022-08-10T11:59:25.670000",
      "content": "<p>Thank you for the great answer I will try the openTsne, it seems a great way to find new antibiotics even for monkeypox.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1900433,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-16T03:18:50.413000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1889671": "This is similar to what I already discussed [**here**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/339726). This time we have a little bit better stacking neural network, and a movie showing the whole process of how data embedding happens with t-SNE. I think it is a good approximation for how the neural network makes the final classification step. Shown at the bottom is the movie's final frame if you want to inspect it at higher resolution.\n\n![t-SNE learning](https://i.ibb.co/MMbDJqV/plot-t-SNE-anim-01.gif)\n\n![t-SNE plot](https://i.ibb.co/Vm9qgvG/t-SNE-keras-training-anim-01.png)\n",
    "1903290": "WOW, So beautiful",
    "1898350": "That is so cool!",
    "1896862": "that's really cool!",
    "1895468": "@tilii7 Fantastic visual & frames of manifold learning.. thanks for sharing!",
    "1893457": "That visualization is incredible, do you mind sharing a sample code, it would be interesting to see how data is being classified in other cases. Does the visualization take long to run on say a sample set of 100,000 rows with 15 features?  ",
    "1892175": "This is really interesting, check my new article on Devgenius classifying bacteria DNA segments using tsne : \nhttps://medium.com/dev-genius/data-science-exploratory-analysis-code-blocks-used-by-data-engineers-ce4529b70672",
    "1889801": "so cool !!",
    "1892020": "The animation is really cool! \n\n\n",
    "1904167": "This is too cool ",
    "1892884": "Thank you for the great answer I will try the openTsne, it seems a great way to find new antibiotics even for monkeypox.",
    "1900433": ""
  }
}