{
  "id": 546584,
  "title": "How about dealing with missing values using an autoencoder?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546584",
  "author_name": "KAI-YI SUNG",
  "post_date": "2024-11-16T17:22:48.510000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Not sure whether this have been discussed but I got an acceptable result by implementing autoencoder. Viellicht this is an acceptable way to handle missing values this time?  Following is my code:</p>\n<p><a href=\"https://www.kaggle.com/code/kaiyisung/missing-value-with-autoencoder\" target=\"_blank\">https://www.kaggle.com/code/kaiyisung/missing-value-with-autoencoder</a></p>\n<p>How did everyone deal with missing values? Please teach me!!</p>",
  "messages": [
    {
      "id": 3047444,
      "postDate": "2024-11-16T17:22:48.510Z",
      "content": "<p>Not sure whether this have been discussed but I got an acceptable result by implementing autoencoder. Viellicht this is an acceptable way to handle missing values this time?  Following is my code:</p>\n<p><a href=\"https://www.kaggle.com/code/kaiyisung/missing-value-with-autoencoder\" target=\"_blank\">https://www.kaggle.com/code/kaiyisung/missing-value-with-autoencoder</a></p>\n<p>How did everyone deal with missing values? Please teach me!!</p>",
      "rawMarkdown": "Not sure whether this have been discussed but I got an acceptable result by implementing autoencoder. Viellicht this is an acceptable way to handle missing values this time?  Following is my code:\n\nhttps://www.kaggle.com/code/kaiyisung/missing-value-with-autoencoder\n\nHow did everyone deal with missing values? Please teach me!!",
      "votes": 3
    },
    {
      "id": 3047790,
      "postDate": "2024-11-17T07:39:06.147Z",
      "content": "<p>Interesting, thanks for sharing! If I understood correctly, wouldn't the reconstruction just return a -1 for the null values?</p>",
      "rawMarkdown": "Interesting, thanks for sharing! If I understood correctly, wouldn't the reconstruction just return a -1 for the null values?",
      "votes": 2,
      "replies": [
        {
          "id": 3047894,
          "postDate": "2024-11-17T10:09:05.300Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/diegoiglesias\" target=\"_blank\">@diegoiglesias</a>! Since the autoencoder cannot take null values as input,  the framework would first replace null values with -1, and hopefully it will learn something during the encoding-decoding process. After that, I filtered out the null values(by selecting -1), and replaced these elements with the result from my autoencoder.</p>\n<p><strong>Potential problems</strong>   </p>\n<ol>\n<li>We lost a huge percentage of data in multiple demisions, even in the target class. At least for me, I don't really believe that there is any good way to deal with this amount of missing data www</li>\n<li>As you noticed, I used -1 to mark null values, which already given the framework a bias. In theory, using 0 is better, but we also have 0 in some columns(especially in sii class)…</li>\n<li>It producted an untrustable result. The result contains around 20% of negative values in sii class. </li>\n</ol>\n<p><strong>What can do next</strong>  </p>\n<ol>\n<li>If the target values are missing, throw them! Perhaps this would help autoencoder work better, there is no way it can predict the target values correctly…I know what my work capable of ;P  </li>\n<li>Still thinking…</li>\n</ol>\n<p>This autoencoder is just for fun!! I'm just a little potato in this field www<br>\ngood luck!</p>",
          "rawMarkdown": "Hi @diegoiglesias! Since the autoencoder cannot take null values as input,  the framework would first replace null values with -1, and hopefully it will learn something during the encoding-decoding process. After that, I filtered out the null values(by selecting -1), and replaced these elements with the result from my autoencoder.\n\n\n**Potential problems**   \n1. We lost a huge percentage of data in multiple demisions, even in the target class. At least for me, I don't really believe that there is any good way to deal with this amount of missing data www\n2. As you noticed, I used -1 to mark null values, which already given the framework a bias. In theory, using 0 is better, but we also have 0 in some columns(especially in sii class)...\n3. It producted an untrustable result. The result contains around 20% of negative values in sii class. \n\n\n**What can do next**  \n1. If the target values are missing, throw them! Perhaps this would help autoencoder work better, there is no way it can predict the target values correctly...I know what my work capable of ;P  \n2. Still thinking...\n\nThis autoencoder is just for fun!! I'm just a little potato in this field www\ngood luck!",
          "votes": 3,
          "replies": [
            {
              "id": 3048284,
              "postDate": "2024-11-17T17:55:40.283Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kaiyisung\" target=\"_blank\">@kaiyisung</a>~</p>\n<p>I have done work related to autoencoder. Based on your description, I thought of two methods:</p>\n<ol>\n<li>Based on the null value problem you mentioned, the attention mechanism can solve this problem well (many public codes currently use <strong>TabNet</strong>, which has a <strong>sparse attention mechanism</strong> that allows the model to focus on only a small part of the input);</li>\n<li>Use the <strong>masking token</strong> from large language models (first proposed in BERT <a href=\"https://arxiv.org/pdf/1810.04805\" target=\"_blank\">here</a>). Specifically, use masking tokens to represent missing data, allowing the autoencoder to learn from the valid information.</li>\n</ol>\n<p>My research is related to LLMs, so I thought of this. If you are interested, you may try it 🖐️</p>",
              "rawMarkdown": "Hi @kaiyisung~\n\nI have done work related to autoencoder. Based on your description, I thought of two methods:\n\n1. Based on the null value problem you mentioned, the attention mechanism can solve this problem well (many public codes currently use **TabNet**, which has a **sparse attention mechanism** that allows the model to focus on only a small part of the input);\n2. Use the **masking token** from large language models (first proposed in BERT [here](https://arxiv.org/pdf/1810.04805)). Specifically, use masking tokens to represent missing data, allowing the autoencoder to learn from the valid information.\n\nMy research is related to LLMs, so I thought of this. If you are interested, you may try it 🖐️",
              "votes": 1
            },
            {
              "id": 3048380,
              "postDate": "2024-11-17T20:19:07.033Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/fangzitao\" target=\"_blank\">@fangzitao</a>! I didn't know I can use masking token in this case, I'll try it!!</p>\n<p>Very out of topic, I'm doing Computational Linguistics and just step in NLP field this semester. Do you have any advice on what should I focus on? Or which tech skills the industry care the most these days? The more I look into this field, the more I realize how weak I am;P</p>",
              "rawMarkdown": "Thanks @fangzitao! I didn't know I can use masking token in this case, I'll try it!!\n\nVery out of topic, I'm doing Computational Linguistics and just step in NLP field this semester. Do you have any advice on what should I focus on? Or which tech skills the industry care the most these days? The more I look into this field, the more I realize how weak I am;P",
              "votes": 1
            },
            {
              "id": 3048614,
              "postDate": "2024-11-18T06:56:57.230Z",
              "content": "<p>Glad to meet you, <a href=\"https://www.kaggle.com/kaiyisung\" target=\"_blank\">@kaiyisung</a>~ Looking forward to what results you can get!</p>\n<p>I'm an undergraduate student majoring in AI and I just share the perspective I have gained through my studies. Hope we can inspire and learn from each other:</p>\n<ul>\n<li>Nowadays, Transformer architecture with the self-attention mechanism \"dominates the AI\". It is a topic you cannot avoid when you come into the AI world, so it is necessary to understand its content;</li>\n<li>Of course, it is built on a lot of basic knowledge of AI and NLP, such as embedding, layer normalization, softmax, and feed-forward networks… So diverse things will appear in the process of learning (🤣 hope you enjoy them);</li>\n<li>If you wanna dive into the research area, both basic and cutting-edge knowledge are necessary. You will definitely find the direction you are interested in during the learning process. Cutting-edge topics are changing all the time, I think it’s more important to find your interests, which is the motivation to keep going;</li>\n<li>For industry, people pay more attention to cost-effectiveness and output. Therefore, some traditional tasks use models with average performance but more stability and strong interpretability instead of DL models like Transformers, depending on the industry. For those large model companies, there are many directions, such as:  <ul>\n<li>LLMs Pre-training / Alignment / CoT (with RDL…)</li>\n<li>SFT to special fields (PEFT…)</li>\n<li>Model Compression &amp; Merging</li>\n<li>Multimodal LMs</li>\n<li>Model quantization/edge deployment</li>\n<li>Transformer-specific optimization/distributed architecture (flash-attn, deepspeed…)</li>\n<li>A lot… </li></ul></li>\n</ul>\n<p>These are just what I know, for your reference. You will find more directions in the process of learning and finally clarify your interests. Enjoy your knowledge trip and happy Kaggling!</p>",
              "rawMarkdown": "Glad to meet you, @kaiyisung~ Looking forward to what results you can get!\n\nI'm an undergraduate student majoring in AI and I just share the perspective I have gained through my studies. Hope we can inspire and learn from each other:\n\n- Nowadays, Transformer architecture with the self-attention mechanism \"dominates the AI\". It is a topic you cannot avoid when you come into the AI world, so it is necessary to understand its content;\n- Of course, it is built on a lot of basic knowledge of AI and NLP, such as embedding, layer normalization, softmax, and feed-forward networks... So diverse things will appear in the process of learning (🤣 hope you enjoy them);\n- If you wanna dive into the research area, both basic and cutting-edge knowledge are necessary. You will definitely find the direction you are interested in during the learning process. Cutting-edge topics are changing all the time, I think it’s more important to find your interests, which is the motivation to keep going;\n- For industry, people pay more attention to cost-effectiveness and output. Therefore, some traditional tasks use models with average performance but more stability and strong interpretability instead of DL models like Transformers, depending on the industry. For those large model companies, there are many directions, such as:  \n  - LLMs Pre-training / Alignment / CoT (with RDL...)\n  - SFT to special fields (PEFT...)\n  - Model Compression & Merging\n  - Multimodal LMs\n  - Model quantization/edge deployment\n  - Transformer-specific optimization/distributed architecture (flash-attn, deepspeed...)\n  - A lot... \n\nThese are just what I know, for your reference. You will find more directions in the process of learning and finally clarify your interests. Enjoy your knowledge trip and happy Kaggling!"
            }
          ]
        }
      ]
    },
    {
      "id": 3047780,
      "postDate": "2024-11-17T07:16:33.267Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3047790,
      "author_name": "dib",
      "author_url": "",
      "post_date": "2024-11-17T07:39:06.147000",
      "content": "<p>Interesting, thanks for sharing! If I understood correctly, wouldn't the reconstruction just return a -1 for the null values?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3047894,
          "author_name": "KAI-YI SUNG",
          "author_url": "",
          "post_date": "2024-11-17T10:09:05.300000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/diegoiglesias\" target=\"_blank\">@diegoiglesias</a>! Since the autoencoder cannot take null values as input,  the framework would first replace null values with -1, and hopefully it will learn something during the encoding-decoding process. After that, I filtered out the null values(by selecting -1), and replaced these elements with the result from my autoencoder.</p>\n<p><strong>Potential problems</strong>   </p>\n<ol>\n<li>We lost a huge percentage of data in multiple demisions, even in the target class. At least for me, I don't really believe that there is any good way to deal with this amount of missing data www</li>\n<li>As you noticed, I used -1 to mark null values, which already given the framework a bias. In theory, using 0 is better, but we also have 0 in some columns(especially in sii class)…</li>\n<li>It producted an untrustable result. The result contains around 20% of negative values in sii class. </li>\n</ol>\n<p><strong>What can do next</strong>  </p>\n<ol>\n<li>If the target values are missing, throw them! Perhaps this would help autoencoder work better, there is no way it can predict the target values correctly…I know what my work capable of ;P  </li>\n<li>Still thinking…</li>\n</ol>\n<p>This autoencoder is just for fun!! I'm just a little potato in this field www<br>\ngood luck!</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3048284,
              "author_name": "Fang Zitao",
              "author_url": "",
              "post_date": "2024-11-17T17:55:40.283000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kaiyisung\" target=\"_blank\">@kaiyisung</a>~</p>\n<p>I have done work related to autoencoder. Based on your description, I thought of two methods:</p>\n<ol>\n<li>Based on the null value problem you mentioned, the attention mechanism can solve this problem well (many public codes currently use <strong>TabNet</strong>, which has a <strong>sparse attention mechanism</strong> that allows the model to focus on only a small part of the input);</li>\n<li>Use the <strong>masking token</strong> from large language models (first proposed in BERT <a href=\"https://arxiv.org/pdf/1810.04805\" target=\"_blank\">here</a>). Specifically, use masking tokens to represent missing data, allowing the autoencoder to learn from the valid information.</li>\n</ol>\n<p>My research is related to LLMs, so I thought of this. If you are interested, you may try it 🖐️</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3048380,
              "author_name": "KAI-YI SUNG",
              "author_url": "",
              "post_date": "2024-11-17T20:19:07.033000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/fangzitao\" target=\"_blank\">@fangzitao</a>! I didn't know I can use masking token in this case, I'll try it!!</p>\n<p>Very out of topic, I'm doing Computational Linguistics and just step in NLP field this semester. Do you have any advice on what should I focus on? Or which tech skills the industry care the most these days? The more I look into this field, the more I realize how weak I am;P</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3048614,
              "author_name": "Fang Zitao",
              "author_url": "",
              "post_date": "2024-11-18T06:56:57.230000",
              "content": "<p>Glad to meet you, <a href=\"https://www.kaggle.com/kaiyisung\" target=\"_blank\">@kaiyisung</a>~ Looking forward to what results you can get!</p>\n<p>I'm an undergraduate student majoring in AI and I just share the perspective I have gained through my studies. Hope we can inspire and learn from each other:</p>\n<ul>\n<li>Nowadays, Transformer architecture with the self-attention mechanism \"dominates the AI\". It is a topic you cannot avoid when you come into the AI world, so it is necessary to understand its content;</li>\n<li>Of course, it is built on a lot of basic knowledge of AI and NLP, such as embedding, layer normalization, softmax, and feed-forward networks… So diverse things will appear in the process of learning (🤣 hope you enjoy them);</li>\n<li>If you wanna dive into the research area, both basic and cutting-edge knowledge are necessary. You will definitely find the direction you are interested in during the learning process. Cutting-edge topics are changing all the time, I think it’s more important to find your interests, which is the motivation to keep going;</li>\n<li>For industry, people pay more attention to cost-effectiveness and output. Therefore, some traditional tasks use models with average performance but more stability and strong interpretability instead of DL models like Transformers, depending on the industry. For those large model companies, there are many directions, such as:  <ul>\n<li>LLMs Pre-training / Alignment / CoT (with RDL…)</li>\n<li>SFT to special fields (PEFT…)</li>\n<li>Model Compression &amp; Merging</li>\n<li>Multimodal LMs</li>\n<li>Model quantization/edge deployment</li>\n<li>Transformer-specific optimization/distributed architecture (flash-attn, deepspeed…)</li>\n<li>A lot… </li></ul></li>\n</ul>\n<p>These are just what I know, for your reference. You will find more directions in the process of learning and finally clarify your interests. Enjoy your knowledge trip and happy Kaggling!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3047780,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-17T07:16:33.267000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3047444": "Not sure whether this have been discussed but I got an acceptable result by implementing autoencoder. Viellicht this is an acceptable way to handle missing values this time?  Following is my code:\n\nhttps://www.kaggle.com/code/kaiyisung/missing-value-with-autoencoder\n\nHow did everyone deal with missing values? Please teach me!!",
    "3047790": "Interesting, thanks for sharing! If I understood correctly, wouldn't the reconstruction just return a -1 for the null values?",
    "3047780": ""
  }
}