{
  "id": 332018,
  "title": "Tabular Deep Learning - A Tutorial",
  "url": "/competitions/amex-default-prediction/discussion/332018",
  "author_name": "",
  "post_date": "2022-06-20T00:50:50.855393900Z",
  "votes": 66,
  "comment_count": 11,
  "views": 0,
  "content": "<h2>Tabular Deep Learning - A Tutorial</h2>\n<p>I had been playing with tabular data &amp; deep learning for the past 2 years as a part of my work and have tried building a plethora of models to tackle different problems: When it comes to deep learning: It is hard. <br>\nLet's review all the current approaches of tabular deep learning you might want to try.</p>\n<hr>\n<h2>Deep Learning &amp; Tabular Data</h2>\n<h3>How tabular data differs from vision &amp; NLP</h3>\n<p><img src=\"https://i.ibb.co/RG8C6jy/spatial-modelling7-24e0998c1f.png\" alt=\"\"></p>\n<p>[<a href=\"https://vsni.co.uk/blogs/spatial_modelling\" target=\"_blank\">source</a>]</p>\n<p>A basic vision model would take an image as input and output an image, the same applies to NLP, taking a sentence as input and outputting the same. There are more detailed considerations such as size, resolution, etc but they are conceptually similar.</p>\n<h5>About Spatial Correlation</h5>\n<p>Images and text are different then tabular data simply because they both have spatial correlation: A small part of an image is correlated to another small part of an image. The same applies to text, where a word is more likely to be in a specific position than another word.</p>\n<p><strong>Intuition: You can simply shuffle the order of the columns in tabular data but you CAN NOT shuffle the order of words in a sentence and expect to end up with a sentence that has the same meaning.</strong></p>\n<p>Tabular data doesn't have spatial correlation, so no part of the data is more likely to be associated with another part, so all the features are important and they are all treated as a whole.</p>\n<h5>Feature Importance</h5>\n<p>Each feature in a tabular dataset is of utmost importance, whereas in vision or NLP we might consider one area of an image as more important than the other, in tabular data, each column is important.</p>\n<h3>Why tabular data has been a pain</h3>\n<p>In the past decade, the amount of experimentation that went on to improve deep learning performance in vision &amp; NLP far surpasses what happened in tabular deep learning. There are many reasons for this but this is out of the scope of this post.</p>\n<p><strong>So, why was it a pain?</strong></p>\n<ul>\n<li>It doesn't really work as well as GBMs. (Most of the time)</li>\n<li>Also: computational resources: Since tabular deep learning often deals with huge fully connected layers we often end up with a network that is very heavy computationally to train.  </li>\n</ul>\n<h2>Tabular Deep Learning Architectures</h2>\n<p>This is a short list of architectures that you can use for tabular data.</p>\n<h3>Simple Fully Connected</h3>\n<p><img src=\"https://i.ibb.co/c2pDMPj/1-gg-Qkn-JWzig-CClfo-pz7-YIA.jpg\" alt=\"\"><br>\n[source<a href=\"https://towardsdatascience.com/tabular-data-analysis-with-deep-neural-nets-d39e10efb6e0\" target=\"_blank\"></a>]</p>\n<p>A simple fully connected network can be thought of as a tabular neural network, where each layer is a fully connected layer, so when we take the input from the input layer, it is multiplied by weights in order to get the output, which then goes to the next layer, etc.</p>\n<p>This is one of the easiest approaches for tabular data: You don't have to worry about padding or embeddings or anything, you just have to care about getting the architecture right. A lot of papers &amp; papers use a simple fully connected network with some feature engineering.</p>\n<p>You can also use a simple FC with attention.</p>\n<h3>CNN</h3>\n<p><img src=\"https://i.ibb.co/jw9kJpt/image-4.png\" alt=\"\"></p>\n<p>A Convolutional Neural Network can be used for tabular data. A simple way to think about a CNN is a simple FC, where you break the input matrix into chunks and slide the window in order to predict each segment.</p>\n<p>The difference is that instead of sliding the window, in a CNN you apply a kernel that goes over the matrix and you have a feature map that you do max pooling on to get the features.</p>\n<p>To use a CNN, you can use a 1D CNN with the feature maps, which is the result of the convolution with the max pooling.</p>\n<p>A CNN with attention to the output of the CNN can give you the most important features.</p>\n<p>You can then use a CNN &amp; max pooling to get the result:</p>\n<ol>\n<li>Use a 1D CNN with a kernel size as the size of your input matrix</li>\n<li>You can also use a sliding window of size n, so you will have n feature maps, then you can max pool them to get the result.</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256\" target=\"_blank\">code</a></p>\n<h3>DeepInsight</h3>\n<p>The idea is very straightforward: Instead of doing feature extraction and selection for collected samples (N samples x d features), we would like to find a way of arranging similar or correlated features into the neighboring regions of a 2-dimensional feature map (d features x N samples) to ease the learning of their complex relationships and interactions. With this general approach, in theory, we could transform any kind of non-image data into feature map images as a friendly representation of samples to CNNs, which provide several unique benefits compared with other neural network architectures, such as automated feature extraction from raw features and memory-footprint reduction by effective weight-sharing.</p>\n<p>The following diagram outlines the key steps. First of all, a non-linear dimensionality reduction technique, like t-SNE or Kernel PCA, is applied to transform raw features into a 2D embeddings feature space. Secondly, the convex hull algorithm is used to find the smallest rectangle containing all features, and a rotation is performed to align the feature map frame into a horizontal or vertical form. Finally, the raw feature values are mapped into the pixel coordinate locations of the feature map image. Note that the resolution of the feature map image affects the ratio of feature overlaps (the features mapped to the same location are averaged), which is a trade-off between the level of lossy compression and computing resource requirements (e.g., host/GPU memory, storage).</p>\n<p><img src=\"https://i.ibb.co/SrQMMyB/deepinsight-architecture-1.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/markpeng/deepinsight-transforming-non-image-data-to-images/notebook\" target=\"_blank\">code</a></p>\n<h3>Tabular Convolution</h3>\n<p><img src=\"https://i.ibb.co/QMxfyWD/cnn-tabular.png\" alt=\"\"></p>\n<p>Tabular convolution is an innovative approach is used to create an image from a tabular sample.</p>\n<ul>\n<li>Choose a sample image (you can consider this as a seed and experiment with various images)<br>\nArrange the input row/sample as a kernel.</li>\n<li>Run a Conv2D on the sample image using this kernel. (In this notebook, I do this operation within the PyTorch model itself)</li>\n<li>Use the resulting image as a sample and use it in your vision model.</li>\n</ul>\n<p>It has started showing promising results but is not yet close to top results. </p>\n<p><a href=\"https://www.kaggle.com/code/krisho007/moa-tabconvolution-training/notebook\" target=\"_blank\">code</a></p>\n<h3>TabNet</h3>\n<p><img src=\"https://i.ibb.co/ccfRVJQ/1-DKm0iz5f-ATCh-DHjl-Wq-NVQ.png\" alt=\"\"></p>\n<p>The first <strong>\"real\"</strong> architecture for tabular data is TabNet, which uses a powerful attention mechanism to get the most important features from the dataset.</p>\n<p>TabNet can be thought of as a TabNet that is applied to the input of a simple fully connected network.</p>\n<p>The attention mechanism is built on the two following steps:</p>\n<ol>\n<li>Encoder: An autoencoder that builds an embedding for each input feature. The embedding is used to weight the feature's importance.</li>\n<li>Decoder: It is just a simple FC network.</li>\n</ol>\n<p>If you have embeddings, it is the same thing, you will not have to worry about the autoencoder.</p>\n<p><a href=\"https://www.kaggle.com/code/tanulsingh077/achieving-sota-results-with-tabnet/notebook\" target=\"_blank\">code</a></p>\n<h3>4. Tab Transformer</h3>\n<p><img src=\"https://i.ibb.co/LnYkCTt/0-d-Rs-H6if2-NC-XLq4.png\" alt=\"\"></p>\n<p>The Tab Transformer is based on the NLP transformer architecture and has similar properties as TabNet.</p>\n<p>A transformer works like this:</p>\n<ol>\n<li>Build an embedding using an autoencoder or using an existing embedding.</li>\n<li>Encoder: An encoder that creates a representation of the input using self-attention</li>\n<li>Decoder: A decoder that takes the representation and produces a result</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/code/cascadinglight/tabtransformer-gauss-rank-baseline-kfold/notebook\" target=\"_blank\">code</a></p>",
  "messages": [
    {
      "id": "1826031",
      "postDate": "06/20/2022 00:50:50",
      "content": "<h2>Tabular Deep Learning - A Tutorial</h2>\n<p>I had been playing with tabular data &amp; deep learning for the past 2 years as a part of my work and have tried building a plethora of models to tackle different problems: When it comes to deep learning: It is hard. <br>\nLet's review all the current approaches of tabular deep learning you might want to try.</p>\n<hr>\n<h2>Deep Learning &amp; Tabular Data</h2>\n<h3>How tabular data differs from vision &amp; NLP</h3>\n<p><img src=\"https://i.ibb.co/RG8C6jy/spatial-modelling7-24e0998c1f.png\" alt=\"\"></p>\n<p>[<a href=\"https://vsni.co.uk/blogs/spatial_modelling\" target=\"_blank\">source</a>]</p>\n<p>A basic vision model would take an image as input and output an image, the same applies to NLP, taking a sentence as input and outputting the same. There are more detailed considerations such as size, resolution, etc but they are conceptually similar.</p>\n<h5>About Spatial Correlation</h5>\n<p>Images and text are different then tabular data simply because they both have spatial correlation: A small part of an image is correlated to another small part of an image. The same applies to text, where a word is more likely to be in a specific position than another word.</p>\n<p><strong>Intuition: You can simply shuffle the order of the columns in tabular data but you CAN NOT shuffle the order of words in a sentence and expect to end up with a sentence that has the same meaning.</strong></p>\n<p>Tabular data doesn't have spatial correlation, so no part of the data is more likely to be associated with another part, so all the features are important and they are all treated as a whole.</p>\n<h5>Feature Importance</h5>\n<p>Each feature in a tabular dataset is of utmost importance, whereas in vision or NLP we might consider one area of an image as more important than the other, in tabular data, each column is important.</p>\n<h3>Why tabular data has been a pain</h3>\n<p>In the past decade, the amount of experimentation that went on to improve deep learning performance in vision &amp; NLP far surpasses what happened in tabular deep learning. There are many reasons for this but this is out of the scope of this post.</p>\n<p><strong>So, why was it a pain?</strong></p>\n<ul>\n<li>It doesn't really work as well as GBMs. (Most of the time)</li>\n<li>Also: computational resources: Since tabular deep learning often deals with huge fully connected layers we often end up with a network that is very heavy computationally to train.  </li>\n</ul>\n<h2>Tabular Deep Learning Architectures</h2>\n<p>This is a short list of architectures that you can use for tabular data.</p>\n<h3>Simple Fully Connected</h3>\n<p><img src=\"https://i.ibb.co/c2pDMPj/1-gg-Qkn-JWzig-CClfo-pz7-YIA.jpg\" alt=\"\"><br>\n[source<a href=\"https://towardsdatascience.com/tabular-data-analysis-with-deep-neural-nets-d39e10efb6e0\" target=\"_blank\"></a>]</p>\n<p>A simple fully connected network can be thought of as a tabular neural network, where each layer is a fully connected layer, so when we take the input from the input layer, it is multiplied by weights in order to get the output, which then goes to the next layer, etc.</p>\n<p>This is one of the easiest approaches for tabular data: You don't have to worry about padding or embeddings or anything, you just have to care about getting the architecture right. A lot of papers &amp; papers use a simple fully connected network with some feature engineering.</p>\n<p>You can also use a simple FC with attention.</p>\n<h3>CNN</h3>\n<p><img src=\"https://i.ibb.co/jw9kJpt/image-4.png\" alt=\"\"></p>\n<p>A Convolutional Neural Network can be used for tabular data. A simple way to think about a CNN is a simple FC, where you break the input matrix into chunks and slide the window in order to predict each segment.</p>\n<p>The difference is that instead of sliding the window, in a CNN you apply a kernel that goes over the matrix and you have a feature map that you do max pooling on to get the features.</p>\n<p>To use a CNN, you can use a 1D CNN with the feature maps, which is the result of the convolution with the max pooling.</p>\n<p>A CNN with attention to the output of the CNN can give you the most important features.</p>\n<p>You can then use a CNN &amp; max pooling to get the result:</p>\n<ol>\n<li>Use a 1D CNN with a kernel size as the size of your input matrix</li>\n<li>You can also use a sliding window of size n, so you will have n feature maps, then you can max pool them to get the result.</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/competitions/lish-moa/discussion/202256\" target=\"_blank\">code</a></p>\n<h3>DeepInsight</h3>\n<p>The idea is very straightforward: Instead of doing feature extraction and selection for collected samples (N samples x d features), we would like to find a way of arranging similar or correlated features into the neighboring regions of a 2-dimensional feature map (d features x N samples) to ease the learning of their complex relationships and interactions. With this general approach, in theory, we could transform any kind of non-image data into feature map images as a friendly representation of samples to CNNs, which provide several unique benefits compared with other neural network architectures, such as automated feature extraction from raw features and memory-footprint reduction by effective weight-sharing.</p>\n<p>The following diagram outlines the key steps. First of all, a non-linear dimensionality reduction technique, like t-SNE or Kernel PCA, is applied to transform raw features into a 2D embeddings feature space. Secondly, the convex hull algorithm is used to find the smallest rectangle containing all features, and a rotation is performed to align the feature map frame into a horizontal or vertical form. Finally, the raw feature values are mapped into the pixel coordinate locations of the feature map image. Note that the resolution of the feature map image affects the ratio of feature overlaps (the features mapped to the same location are averaged), which is a trade-off between the level of lossy compression and computing resource requirements (e.g., host/GPU memory, storage).</p>\n<p><img src=\"https://i.ibb.co/SrQMMyB/deepinsight-architecture-1.png\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/markpeng/deepinsight-transforming-non-image-data-to-images/notebook\" target=\"_blank\">code</a></p>\n<h3>Tabular Convolution</h3>\n<p><img src=\"https://i.ibb.co/QMxfyWD/cnn-tabular.png\" alt=\"\"></p>\n<p>Tabular convolution is an innovative approach is used to create an image from a tabular sample.</p>\n<ul>\n<li>Choose a sample image (you can consider this as a seed and experiment with various images)<br>\nArrange the input row/sample as a kernel.</li>\n<li>Run a Conv2D on the sample image using this kernel. (In this notebook, I do this operation within the PyTorch model itself)</li>\n<li>Use the resulting image as a sample and use it in your vision model.</li>\n</ul>\n<p>It has started showing promising results but is not yet close to top results. </p>\n<p><a href=\"https://www.kaggle.com/code/krisho007/moa-tabconvolution-training/notebook\" target=\"_blank\">code</a></p>\n<h3>TabNet</h3>\n<p><img src=\"https://i.ibb.co/ccfRVJQ/1-DKm0iz5f-ATCh-DHjl-Wq-NVQ.png\" alt=\"\"></p>\n<p>The first <strong>\"real\"</strong> architecture for tabular data is TabNet, which uses a powerful attention mechanism to get the most important features from the dataset.</p>\n<p>TabNet can be thought of as a TabNet that is applied to the input of a simple fully connected network.</p>\n<p>The attention mechanism is built on the two following steps:</p>\n<ol>\n<li>Encoder: An autoencoder that builds an embedding for each input feature. The embedding is used to weight the feature's importance.</li>\n<li>Decoder: It is just a simple FC network.</li>\n</ol>\n<p>If you have embeddings, it is the same thing, you will not have to worry about the autoencoder.</p>\n<p><a href=\"https://www.kaggle.com/code/tanulsingh077/achieving-sota-results-with-tabnet/notebook\" target=\"_blank\">code</a></p>\n<h3>4. Tab Transformer</h3>\n<p><img src=\"https://i.ibb.co/LnYkCTt/0-d-Rs-H6if2-NC-XLq4.png\" alt=\"\"></p>\n<p>The Tab Transformer is based on the NLP transformer architecture and has similar properties as TabNet.</p>\n<p>A transformer works like this:</p>\n<ol>\n<li>Build an embedding using an autoencoder or using an existing embedding.</li>\n<li>Encoder: An encoder that creates a representation of the input using self-attention</li>\n<li>Decoder: A decoder that takes the representation and produces a result</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/code/cascadinglight/tabtransformer-gauss-rank-baseline-kfold/notebook\" target=\"_blank\">code</a></p>",
      "rawMarkdown": "## Tabular Deep Learning - A Tutorial\n\nI had been playing with tabular data & deep learning for the past 2 years as a part of my work and have tried building a plethora of models to tackle different problems: When it comes to deep learning: It is hard. \nLet's review all the current approaches of tabular deep learning you might want to try.\n\n_____\n\n## Deep Learning & Tabular Data\n### How tabular data differs from vision & NLP\n\n![](https://i.ibb.co/RG8C6jy/spatial-modelling7-24e0998c1f.png)\n\n[[source](https://vsni.co.uk/blogs/spatial_modelling)]\n\nA basic vision model would take an image as input and output an image, the same applies to NLP, taking a sentence as input and outputting the same. There are more detailed considerations such as size, resolution, etc but they are conceptually similar.\n\n##### About Spatial Correlation\nImages and text are different then tabular data simply because they both have spatial correlation: A small part of an image is correlated to another small part of an image. The same applies to text, where a word is more likely to be in a specific position than another word.\n\n**Intuition: You can simply shuffle the order of the columns in tabular data but you CAN NOT shuffle the order of words in a sentence and expect to end up with a sentence that has the same meaning.**\n\nTabular data doesn't have spatial correlation, so no part of the data is more likely to be associated with another part, so all the features are important and they are all treated as a whole.\n\n##### Feature Importance\nEach feature in a tabular dataset is of utmost importance, whereas in vision or NLP we might consider one area of an image as more important than the other, in tabular data, each column is important.\n\n### Why tabular data has been a pain\n\nIn the past decade, the amount of experimentation that went on to improve deep learning performance in vision & NLP far surpasses what happened in tabular deep learning. There are many reasons for this but this is out of the scope of this post.\n\n**So, why was it a pain?**\n- It doesn't really work as well as GBMs. (Most of the time)\n- Also: computational resources: Since tabular deep learning often deals with huge fully connected layers we often end up with a network that is very heavy computationally to train.  \n\n## Tabular Deep Learning Architectures\n\nThis is a short list of architectures that you can use for tabular data.\n\n### Simple Fully Connected\n\n![](https://i.ibb.co/c2pDMPj/1-gg-Qkn-JWzig-CClfo-pz7-YIA.jpg)\n[source[](https://towardsdatascience.com/tabular-data-analysis-with-deep-neural-nets-d39e10efb6e0)]\n\nA simple fully connected network can be thought of as a tabular neural network, where each layer is a fully connected layer, so when we take the input from the input layer, it is multiplied by weights in order to get the output, which then goes to the next layer, etc.\n\nThis is one of the easiest approaches for tabular data: You don't have to worry about padding or embeddings or anything, you just have to care about getting the architecture right. A lot of papers & papers use a simple fully connected network with some feature engineering.\n\nYou can also use a simple FC with attention.\n\n### CNN\n\n![](https://i.ibb.co/jw9kJpt/image-4.png)\n\nA Convolutional Neural Network can be used for tabular data. A simple way to think about a CNN is a simple FC, where you break the input matrix into chunks and slide the window in order to predict each segment.\n\nThe difference is that instead of sliding the window, in a CNN you apply a kernel that goes over the matrix and you have a feature map that you do max pooling on to get the features.\n\nTo use a CNN, you can use a 1D CNN with the feature maps, which is the result of the convolution with the max pooling.\n\nA CNN with attention to the output of the CNN can give you the most important features.\n\nYou can then use a CNN & max pooling to get the result:\n\n1. Use a 1D CNN with a kernel size as the size of your input matrix\n2. You can also use a sliding window of size n, so you will have n feature maps, then you can max pool them to get the result.\n\n[code](https://www.kaggle.com/competitions/lish-moa/discussion/202256)\n\n### DeepInsight\n\nThe idea is very straightforward: Instead of doing feature extraction and selection for collected samples (N samples x d features), we would like to find a way of arranging similar or correlated features into the neighboring regions of a 2-dimensional feature map (d features x N samples) to ease the learning of their complex relationships and interactions. With this general approach, in theory, we could transform any kind of non-image data into feature map images as a friendly representation of samples to CNNs, which provide several unique benefits compared with other neural network architectures, such as automated feature extraction from raw features and memory-footprint reduction by effective weight-sharing.\n\nThe following diagram outlines the key steps. First of all, a non-linear dimensionality reduction technique, like t-SNE or Kernel PCA, is applied to transform raw features into a 2D embeddings feature space. Secondly, the convex hull algorithm is used to find the smallest rectangle containing all features, and a rotation is performed to align the feature map frame into a horizontal or vertical form. Finally, the raw feature values are mapped into the pixel coordinate locations of the feature map image. Note that the resolution of the feature map image affects the ratio of feature overlaps (the features mapped to the same location are averaged), which is a trade-off between the level of lossy compression and computing resource requirements (e.g., host/GPU memory, storage).\n\n![](https://i.ibb.co/SrQMMyB/deepinsight-architecture-1.png)\n\n[code](https://www.kaggle.com/code/markpeng/deepinsight-transforming-non-image-data-to-images/notebook)\n\n### Tabular Convolution\n\n![](https://i.ibb.co/QMxfyWD/cnn-tabular.png)\n\n\nTabular convolution is an innovative approach is used to create an image from a tabular sample.\n\n- Choose a sample image (you can consider this as a seed and experiment with various images)\nArrange the input row/sample as a kernel.\n- Run a Conv2D on the sample image using this kernel. (In this notebook, I do this operation within the PyTorch model itself)\n- Use the resulting image as a sample and use it in your vision model.\n\nIt has started showing promising results but is not yet close to top results. \n\n[code](https://www.kaggle.com/code/krisho007/moa-tabconvolution-training/notebook)\n\n### TabNet\n\n![](https://i.ibb.co/ccfRVJQ/1-DKm0iz5f-ATCh-DHjl-Wq-NVQ.png)\n\nThe first **\"real\"** architecture for tabular data is TabNet, which uses a powerful attention mechanism to get the most important features from the dataset.\n\nTabNet can be thought of as a TabNet that is applied to the input of a simple fully connected network.\n\nThe attention mechanism is built on the two following steps:\n\n1. Encoder: An autoencoder that builds an embedding for each input feature. The embedding is used to weight the feature's importance.\n2. Decoder: It is just a simple FC network.\n\nIf you have embeddings, it is the same thing, you will not have to worry about the autoencoder.\n\n[code](https://www.kaggle.com/code/tanulsingh077/achieving-sota-results-with-tabnet/notebook)\n\n### 4. Tab Transformer\n\n![](https://i.ibb.co/LnYkCTt/0-d-Rs-H6if2-NC-XLq4.png)\n\nThe Tab Transformer is based on the NLP transformer architecture and has similar properties as TabNet.\n\nA transformer works like this:\n\n1. Build an embedding using an autoencoder or using an existing embedding.\n2. Encoder: An encoder that creates a representation of the input using self-attention\n3. Decoder: A decoder that takes the representation and produces a result\n\n[code](https://www.kaggle.com/code/cascadinglight/tabtransformer-gauss-rank-baseline-kfold/notebook)",
      "votes": null
    },
    {
      "id": "1826037",
      "postDate": "06/20/2022 01:18:25",
      "content": "<p>Amazing, helpful tutorial and its source too.</p>",
      "rawMarkdown": "Amazing, helpful tutorial and its source too.",
      "votes": null
    },
    {
      "id": "1826051",
      "postDate": "06/20/2022 01:42:09",
      "content": "<p>amazing summary. Thank you for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "rawMarkdown": "amazing summary. Thank you for sharing @thedevastator",
      "votes": null
    },
    {
      "id": "1826225",
      "postDate": "06/20/2022 07:07:26",
      "content": "<p>My tiny mind is blown! Thank you for this!</p>",
      "rawMarkdown": "My tiny mind is blown! Thank you for this!",
      "votes": null
    },
    {
      "id": "1826448",
      "postDate": "06/20/2022 10:57:05",
      "content": "<p>Nice summary!</p>",
      "rawMarkdown": "Nice summary!",
      "votes": null
    },
    {
      "id": "1826565",
      "postDate": "06/20/2022 12:57:19",
      "content": "<p>I need a summary like this. Thanks! </p>",
      "rawMarkdown": "I need a summary like this. Thanks!",
      "votes": null
    },
    {
      "id": "1827465",
      "postDate": "06/21/2022 04:28:46",
      "content": "<p>Nice summary. You forgot to mention Denoising Autoencoders. They have won many previous Kaggle tabular data competitions.</p>",
      "rawMarkdown": "Nice summary. You forgot to mention Denoising Autoencoders. They have won many previous Kaggle tabular data competitions.",
      "votes": null
    },
    {
      "id": "1829529",
      "postDate": "06/22/2022 17:08:11",
      "content": "<p>This is just what I needed to understand this. Amazing work👍😁</p>",
      "rawMarkdown": "This is just what I needed to understand this. Amazing work👍😁",
      "votes": null
    },
    {
      "id": "1834526",
      "postDate": "06/27/2022 02:27:49",
      "content": "<p>Good summarization of the popular techniques , also code references are very helpful . will post my results once I try those currently focusing on xg and lgbm  . Do post your results as well using these approaches .</p>",
      "rawMarkdown": "Good summarization of the popular techniques , also code references are very helpful . will post my results once I try those currently focusing on xg and lgbm  . Do post your results as well using these approaches .",
      "votes": null
    },
    {
      "id": "1852075",
      "postDate": "07/11/2022 18:27:41",
      "content": "<p>This is a great summary. Thanks for the amazing work!</p>",
      "rawMarkdown": "This is a great summary. Thanks for the amazing work!",
      "votes": null
    },
    {
      "id": "1877987",
      "postDate": "07/31/2022 05:32:18",
      "content": "<p>Great synopsis! Learnt a lot. <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> . </p>",
      "rawMarkdown": "Great synopsis! Learnt a lot. @thedevastator .",
      "votes": null
    },
    {
      "id": "1915158",
      "postDate": "08/26/2022 18:02:06",
      "content": "<p>insightful discussion </p>",
      "rawMarkdown": "insightful discussion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1826037,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "06/20/2022 01:18:25",
      "content": "<p>Amazing, helpful tutorial and its source too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1826051,
      "author_name": "naiborhujosua",
      "author_url": "",
      "post_date": "06/20/2022 01:42:09",
      "content": "<p>amazing summary. Thank you for sharing <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1826225,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "06/20/2022 07:07:26",
      "content": "<p>My tiny mind is blown! Thank you for this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1826448,
      "author_name": "paulojunqueira",
      "author_url": "",
      "post_date": "06/20/2022 10:57:05",
      "content": "<p>Nice summary!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1826565,
      "author_name": "kimkeonho",
      "author_url": "",
      "post_date": "06/20/2022 12:57:19",
      "content": "<p>I need a summary like this. Thanks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1827465,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/21/2022 04:28:46",
      "content": "<p>Nice summary. You forgot to mention Denoising Autoencoders. They have won many previous Kaggle tabular data competitions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1829529,
      "author_name": "durgancegaur",
      "author_url": "",
      "post_date": "06/22/2022 17:08:11",
      "content": "<p>This is just what I needed to understand this. Amazing work👍😁</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1834526,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "06/27/2022 02:27:49",
      "content": "<p>Good summarization of the popular techniques , also code references are very helpful . will post my results once I try those currently focusing on xg and lgbm  . Do post your results as well using these approaches .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1852075,
      "author_name": "jnnerd",
      "author_url": "",
      "post_date": "07/11/2022 18:27:41",
      "content": "<p>This is a great summary. Thanks for the amazing work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1877987,
      "author_name": "casper6290",
      "author_url": "",
      "post_date": "07/31/2022 05:32:18",
      "content": "<p>Great synopsis! Learnt a lot. <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> . </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1915158,
      "author_name": "gazu468",
      "author_url": "",
      "post_date": "08/26/2022 18:02:06",
      "content": "<p>insightful discussion </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1826031": "## Tabular Deep Learning - A Tutorial\n\nI had been playing with tabular data & deep learning for the past 2 years as a part of my work and have tried building a plethora of models to tackle different problems: When it comes to deep learning: It is hard. \nLet's review all the current approaches of tabular deep learning you might want to try.\n\n_____\n\n## Deep Learning & Tabular Data\n### How tabular data differs from vision & NLP\n\n![](https://i.ibb.co/RG8C6jy/spatial-modelling7-24e0998c1f.png)\n\n[[source](https://vsni.co.uk/blogs/spatial_modelling)]\n\nA basic vision model would take an image as input and output an image, the same applies to NLP, taking a sentence as input and outputting the same. There are more detailed considerations such as size, resolution, etc but they are conceptually similar.\n\n##### About Spatial Correlation\nImages and text are different then tabular data simply because they both have spatial correlation: A small part of an image is correlated to another small part of an image. The same applies to text, where a word is more likely to be in a specific position than another word.\n\n**Intuition: You can simply shuffle the order of the columns in tabular data but you CAN NOT shuffle the order of words in a sentence and expect to end up with a sentence that has the same meaning.**\n\nTabular data doesn't have spatial correlation, so no part of the data is more likely to be associated with another part, so all the features are important and they are all treated as a whole.\n\n##### Feature Importance\nEach feature in a tabular dataset is of utmost importance, whereas in vision or NLP we might consider one area of an image as more important than the other, in tabular data, each column is important.\n\n### Why tabular data has been a pain\n\nIn the past decade, the amount of experimentation that went on to improve deep learning performance in vision & NLP far surpasses what happened in tabular deep learning. There are many reasons for this but this is out of the scope of this post.\n\n**So, why was it a pain?**\n- It doesn't really work as well as GBMs. (Most of the time)\n- Also: computational resources: Since tabular deep learning often deals with huge fully connected layers we often end up with a network that is very heavy computationally to train.  \n\n## Tabular Deep Learning Architectures\n\nThis is a short list of architectures that you can use for tabular data.\n\n### Simple Fully Connected\n\n![](https://i.ibb.co/c2pDMPj/1-gg-Qkn-JWzig-CClfo-pz7-YIA.jpg)\n[source[](https://towardsdatascience.com/tabular-data-analysis-with-deep-neural-nets-d39e10efb6e0)]\n\nA simple fully connected network can be thought of as a tabular neural network, where each layer is a fully connected layer, so when we take the input from the input layer, it is multiplied by weights in order to get the output, which then goes to the next layer, etc.\n\nThis is one of the easiest approaches for tabular data: You don't have to worry about padding or embeddings or anything, you just have to care about getting the architecture right. A lot of papers & papers use a simple fully connected network with some feature engineering.\n\nYou can also use a simple FC with attention.\n\n### CNN\n\n![](https://i.ibb.co/jw9kJpt/image-4.png)\n\nA Convolutional Neural Network can be used for tabular data. A simple way to think about a CNN is a simple FC, where you break the input matrix into chunks and slide the window in order to predict each segment.\n\nThe difference is that instead of sliding the window, in a CNN you apply a kernel that goes over the matrix and you have a feature map that you do max pooling on to get the features.\n\nTo use a CNN, you can use a 1D CNN with the feature maps, which is the result of the convolution with the max pooling.\n\nA CNN with attention to the output of the CNN can give you the most important features.\n\nYou can then use a CNN & max pooling to get the result:\n\n1. Use a 1D CNN with a kernel size as the size of your input matrix\n2. You can also use a sliding window of size n, so you will have n feature maps, then you can max pool them to get the result.\n\n[code](https://www.kaggle.com/competitions/lish-moa/discussion/202256)\n\n### DeepInsight\n\nThe idea is very straightforward: Instead of doing feature extraction and selection for collected samples (N samples x d features), we would like to find a way of arranging similar or correlated features into the neighboring regions of a 2-dimensional feature map (d features x N samples) to ease the learning of their complex relationships and interactions. With this general approach, in theory, we could transform any kind of non-image data into feature map images as a friendly representation of samples to CNNs, which provide several unique benefits compared with other neural network architectures, such as automated feature extraction from raw features and memory-footprint reduction by effective weight-sharing.\n\nThe following diagram outlines the key steps. First of all, a non-linear dimensionality reduction technique, like t-SNE or Kernel PCA, is applied to transform raw features into a 2D embeddings feature space. Secondly, the convex hull algorithm is used to find the smallest rectangle containing all features, and a rotation is performed to align the feature map frame into a horizontal or vertical form. Finally, the raw feature values are mapped into the pixel coordinate locations of the feature map image. Note that the resolution of the feature map image affects the ratio of feature overlaps (the features mapped to the same location are averaged), which is a trade-off between the level of lossy compression and computing resource requirements (e.g., host/GPU memory, storage).\n\n![](https://i.ibb.co/SrQMMyB/deepinsight-architecture-1.png)\n\n[code](https://www.kaggle.com/code/markpeng/deepinsight-transforming-non-image-data-to-images/notebook)\n\n### Tabular Convolution\n\n![](https://i.ibb.co/QMxfyWD/cnn-tabular.png)\n\n\nTabular convolution is an innovative approach is used to create an image from a tabular sample.\n\n- Choose a sample image (you can consider this as a seed and experiment with various images)\nArrange the input row/sample as a kernel.\n- Run a Conv2D on the sample image using this kernel. (In this notebook, I do this operation within the PyTorch model itself)\n- Use the resulting image as a sample and use it in your vision model.\n\nIt has started showing promising results but is not yet close to top results. \n\n[code](https://www.kaggle.com/code/krisho007/moa-tabconvolution-training/notebook)\n\n### TabNet\n\n![](https://i.ibb.co/ccfRVJQ/1-DKm0iz5f-ATCh-DHjl-Wq-NVQ.png)\n\nThe first **\"real\"** architecture for tabular data is TabNet, which uses a powerful attention mechanism to get the most important features from the dataset.\n\nTabNet can be thought of as a TabNet that is applied to the input of a simple fully connected network.\n\nThe attention mechanism is built on the two following steps:\n\n1. Encoder: An autoencoder that builds an embedding for each input feature. The embedding is used to weight the feature's importance.\n2. Decoder: It is just a simple FC network.\n\nIf you have embeddings, it is the same thing, you will not have to worry about the autoencoder.\n\n[code](https://www.kaggle.com/code/tanulsingh077/achieving-sota-results-with-tabnet/notebook)\n\n### 4. Tab Transformer\n\n![](https://i.ibb.co/LnYkCTt/0-d-Rs-H6if2-NC-XLq4.png)\n\nThe Tab Transformer is based on the NLP transformer architecture and has similar properties as TabNet.\n\nA transformer works like this:\n\n1. Build an embedding using an autoencoder or using an existing embedding.\n2. Encoder: An encoder that creates a representation of the input using self-attention\n3. Decoder: A decoder that takes the representation and produces a result\n\n[code](https://www.kaggle.com/code/cascadinglight/tabtransformer-gauss-rank-baseline-kfold/notebook)",
    "1826037": "Amazing, helpful tutorial and its source too.",
    "1826051": "amazing summary. Thank you for sharing @thedevastator",
    "1826225": "My tiny mind is blown! Thank you for this!",
    "1826448": "Nice summary!",
    "1826565": "I need a summary like this. Thanks!",
    "1827465": "Nice summary. You forgot to mention Denoising Autoencoders. They have won many previous Kaggle tabular data competitions.",
    "1829529": "This is just what I needed to understand this. Amazing work👍😁",
    "1834526": "Good summarization of the popular techniques , also code references are very helpful . will post my results once I try those currently focusing on xg and lgbm  . Do post your results as well using these approaches .",
    "1852075": "This is a great summary. Thanks for the amazing work!",
    "1877987": "Great synopsis! Learnt a lot. @thedevastator .",
    "1915158": "insightful discussion"
  },
  "source": "meta"
}