{
  "id": 347651,
  "title": "19th Place Solution",
  "url": "/competitions/amex-default-prediction/writeups/fritz-cremer-19th-place-solution",
  "author_name": "",
  "post_date": "2022-08-30T09:28:26.730Z",
  "votes": 52,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Thanks to AMEX for hosting this competition!<br>\nI just want to give a quick overview of the stuff I did in this competition.</p>\n<h1>Main concepts</h1>\n<p>I used an NN based models and LGBM based models for this competition. </p>\n<h4>Pre-processing</h4>\n<p>Because of some features with very spread out distributions, I figured that log-transforming several features for the NN models might make sense. I log-transformed a good fraction of the features. I decided to transform or not by gut feeling after looking at the histogram of each feature.</p>\n<p>I noticed 2 features that behaved weirdly in private data, so I removed them for all NN models.<br>\nThose features <strong>D_59</strong>, <strong>D_86</strong>. <strong>D_59</strong> even had a different private test data distribution than public test data distribution.</p>\n<h4>Mini-LSTMs</h4>\n<p>For both model types, I constructed new features by training a simple LSTM on every single feature (13x1 input dimension). Those models still predict the original target.<br>\nThis results in 189 different models which generate probabilities given the sequence of only a single feature.</p>\n<h4>LGBMs</h4>\n<p>For the LGBM based model, I used the great <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> as a starting point. I added the previously generated features from the LSTM models and changed tweaked a few hyperparameters. I also did one training with all data.</p>\n<h4>NN</h4>\n<p>For the NN I have first pre-trained a transformer-model on all data which imputed random missing values. I then embedded this pre-trained model into an NN which also used several other features (e.g. last sequence element) and trained it. This improved the NN performance quite a bit. I also passed the Mini-LSTM model features into the NN.</p>\n<h4>Pseudo-labeling</h4>\n<p>I also trained the previous NN on pseudo labeled test data (labeled by all other previous models), which gave this nn a score of 0.799 on public LB. This was however difficult the validate, since the pseudo labels created a heavy leak. Therefore I validated this model only on the public LB data. This helped to further increase my score.</p>\n<h4>Final model</h4>\n<p>In the end I ensembled the NNs and LGBMs with roughly 30/70 (NN/LGBM) weight.<br>\nI also included this <a href=\"https://www.kaggle.com/code/swimmy/tuffline-amex-anotherfeaturelgbm\" target=\"_blank\">public submission</a> in the ensemble.</p>\n<h1>Things that did not work</h1>\n<p>Or at least things that didn't help much.</p>\n<p>I tried using pre-trained autoencoders, ELECTRA style pretraining, Tabnet, CNN, LSTM, XGB, other losses (pearson, mse), and some smaller stuff…<br>\nI tried <strong>a lot</strong> of things with some of these but it just didn't really help. I actually used embeddings generated by the ELECTRA style pre-trained transformer model for my NN, but I think this only contributes to a minor improvement.</p>\n<p>Anyway, that's basically it. I am sure all of it 100-200 hours and I feel like there is still a lot of room for improvement (e.g. I did not test enough features for LGBMs, could still train more models since I only use like 25 single models in the final version). I focused a lot on playing with NN in this competition, but a lot of stuff just wouldn't work.</p>",
  "messages": [
    {
      "id": "1912797",
      "postDate": "08/25/2022 01:13:36",
      "content": "<p>Thanks to AMEX for hosting this competition!<br>\nI just want to give a quick overview of the stuff I did in this competition.</p>\n<h1>Main concepts</h1>\n<p>I used an NN based models and LGBM based models for this competition. </p>\n<h4>Pre-processing</h4>\n<p>Because of some features with very spread out distributions, I figured that log-transforming several features for the NN models might make sense. I log-transformed a good fraction of the features. I decided to transform or not by gut feeling after looking at the histogram of each feature.</p>\n<p>I noticed 2 features that behaved weirdly in private data, so I removed them for all NN models.<br>\nThose features <strong>D_59</strong>, <strong>D_86</strong>. <strong>D_59</strong> even had a different private test data distribution than public test data distribution.</p>\n<h4>Mini-LSTMs</h4>\n<p>For both model types, I constructed new features by training a simple LSTM on every single feature (13x1 input dimension). Those models still predict the original target.<br>\nThis results in 189 different models which generate probabilities given the sequence of only a single feature.</p>\n<h4>LGBMs</h4>\n<p>For the LGBM based model, I used the great <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> as a starting point. I added the previously generated features from the LSTM models and changed tweaked a few hyperparameters. I also did one training with all data.</p>\n<h4>NN</h4>\n<p>For the NN I have first pre-trained a transformer-model on all data which imputed random missing values. I then embedded this pre-trained model into an NN which also used several other features (e.g. last sequence element) and trained it. This improved the NN performance quite a bit. I also passed the Mini-LSTM model features into the NN.</p>\n<h4>Pseudo-labeling</h4>\n<p>I also trained the previous NN on pseudo labeled test data (labeled by all other previous models), which gave this nn a score of 0.799 on public LB. This was however difficult the validate, since the pseudo labels created a heavy leak. Therefore I validated this model only on the public LB data. This helped to further increase my score.</p>\n<h4>Final model</h4>\n<p>In the end I ensembled the NNs and LGBMs with roughly 30/70 (NN/LGBM) weight.<br>\nI also included this <a href=\"https://www.kaggle.com/code/swimmy/tuffline-amex-anotherfeaturelgbm\" target=\"_blank\">public submission</a> in the ensemble.</p>\n<h1>Things that did not work</h1>\n<p>Or at least things that didn't help much.</p>\n<p>I tried using pre-trained autoencoders, ELECTRA style pretraining, Tabnet, CNN, LSTM, XGB, other losses (pearson, mse), and some smaller stuff…<br>\nI tried <strong>a lot</strong> of things with some of these but it just didn't really help. I actually used embeddings generated by the ELECTRA style pre-trained transformer model for my NN, but I think this only contributes to a minor improvement.</p>\n<p>Anyway, that's basically it. I am sure all of it 100-200 hours and I feel like there is still a lot of room for improvement (e.g. I did not test enough features for LGBMs, could still train more models since I only use like 25 single models in the final version). I focused a lot on playing with NN in this competition, but a lot of stuff just wouldn't work.</p>",
      "rawMarkdown": "Thanks to AMEX for hosting this competition!\nI just want to give a quick overview of the stuff I did in this competition.\n\n# Main concepts\n\nI used an NN based models and LGBM based models for this competition. \n#### Pre-processing \nBecause of some features with very spread out distributions, I figured that log-transforming several features for the NN models might make sense. I log-transformed a good fraction of the features. I decided to transform or not by gut feeling after looking at the histogram of each feature.\n\nI noticed 2 features that behaved weirdly in private data, so I removed them for all NN models.\nThose features **D_59**, **D_86**. **D_59** even had a different private test data distribution than public test data distribution.\n\n#### Mini-LSTMs\nFor both model types, I constructed new features by training a simple LSTM on every single feature (13x1 input dimension). Those models still predict the original target.\nThis results in 189 different models which generate probabilities given the sequence of only a single feature.\n\n#### LGBMs\nFor the LGBM based model, I used the great [notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977) by @ragnar123 as a starting point. I added the previously generated features from the LSTM models and changed tweaked a few hyperparameters. I also did one training with all data.\n\n#### NN\nFor the NN I have first pre-trained a transformer-model on all data which imputed random missing values. I then embedded this pre-trained model into an NN which also used several other features (e.g. last sequence element) and trained it. This improved the NN performance quite a bit. I also passed the Mini-LSTM model features into the NN.\n\n#### Pseudo-labeling\nI also trained the previous NN on pseudo labeled test data (labeled by all other previous models), which gave this nn a score of 0.799 on public LB. This was however difficult the validate, since the pseudo labels created a heavy leak. Therefore I validated this model only on the public LB data. This helped to further increase my score.\n\n#### Final model\nIn the end I ensembled the NNs and LGBMs with roughly 30/70 (NN/LGBM) weight.\nI also included this [public submission](https://www.kaggle.com/code/swimmy/tuffline-amex-anotherfeaturelgbm) in the ensemble.\n\n# Things that did not work\nOr at least things that didn't help much.\n\nI tried using pre-trained autoencoders, ELECTRA style pretraining, Tabnet, CNN, LSTM, XGB, other losses (pearson, mse), and some smaller stuff...\nI tried **a lot** of things with some of these but it just didn't really help. I actually used embeddings generated by the ELECTRA style pre-trained transformer model for my NN, but I think this only contributes to a minor improvement.\n\nAnyway, that's basically it. I am sure all of it 100-200 hours and I feel like there is still a lot of room for improvement (e.g. I did not test enough features for LGBMs, could still train more models since I only use like 25 single models in the final version). I focused a lot on playing with NN in this competition, but a lot of stuff just wouldn't work.",
      "votes": null
    },
    {
      "id": "1912849",
      "postDate": "08/25/2022 02:08:38",
      "content": "<p>Great work! Congrats on your achievement!</p>\n<p>We used concept which is pretty similar to your mini-LSTM, we found that train lightgbm on each feature produce a better results.</p>\n<p>You can find the notebook <a href=\"https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431\" target=\"_blank\">here</a><br>\nAny chance that you will share your solution as a notebook?</p>",
      "rawMarkdown": "Great work! Congrats on your achievement!\n\nWe used concept which is pretty similar to your mini-LSTM, we found that train lightgbm on each feature produce a better results.\n\nYou can find the notebook [here](https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431)\nAny chance that you will share your solution as a notebook?",
      "votes": null
    },
    {
      "id": "1912934",
      "postDate": "08/25/2022 03:55:44",
      "content": "<p>Congratulations and thanks for sharing the solution. Can you please elaborate on the Electra Pretraining part?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing the solution. Can you please elaborate on the Electra Pretraining part?",
      "votes": null
    },
    {
      "id": "1913175",
      "postDate": "08/25/2022 07:44:32",
      "content": "<p>I second the request </p>",
      "rawMarkdown": "I second the request",
      "votes": null
    },
    {
      "id": "1913227",
      "postDate": "08/25/2022 08:38:45",
      "content": "<p>Yes, so I though about a way to generate embeddings by pre-training on all data. I was inspired by the training process of Google's language model ELECTRA. I used a generator and a discriminator. The generator learned to produce sequence elements (e.g. the entire last statement) and swapped them with the actual data. The discriminator needed to decide, which of the 13 sequence elements have been created by the generator (e.g. 5, 7, 10, 11 have been produced by the generator). This produced embeddings in the last hidden layer that had quite a decent predictive power (though way too weak to use on it's own), so I experimented with merging those features with the NN and the LGBM. With the NN I saw a minor improvement and with the LGBM no improvement (this was probably due to noise since each training went a bit different anyways). For the NN it also could have been noise but I just kept it in.</p>\n<p>And actually I did not add the features itself, but rather I trained another NN to produce a single probabiltiy from those embeddings, and fed this single feature to the NN. Overall, my pipeline feels a bit complicated and things like these could probably be removed.</p>",
      "rawMarkdown": "Yes, so I though about a way to generate embeddings by pre-training on all data. I was inspired by the training process of Google's language model ELECTRA. I used a generator and a discriminator. The generator learned to produce sequence elements (e.g. the entire last statement) and swapped them with the actual data. The discriminator needed to decide, which of the 13 sequence elements have been created by the generator (e.g. 5, 7, 10, 11 have been produced by the generator). This produced embeddings in the last hidden layer that had quite a decent predictive power (though way too weak to use on it's own), so I experimented with merging those features with the NN and the LGBM. With the NN I saw a minor improvement and with the LGBM no improvement (this was probably due to noise since each training went a bit different anyways). For the NN it also could have been noise but I just kept it in.\n\nAnd actually I did not add the features itself, but rather I trained another NN to produce a single probabiltiy from those embeddings, and fed this single feature to the NN. Overall, my pipeline feels a bit complicated and things like these could probably be removed.",
      "votes": null
    },
    {
      "id": "1913229",
      "postDate": "08/25/2022 08:41:08",
      "content": "<p>Oh that's interesting, I though about an LGBM approach as well but thought LSTM might outperform them, because of the sequence structure and relatively low dimensionality of the inputs. For sharing, I am not sure how feasible it is. It would be quite a mess since so many files are involved and I worked locally on my computer.</p>",
      "rawMarkdown": "Oh that's interesting, I though about an LGBM approach as well but thought LSTM might outperform them, because of the sequence structure and relatively low dimensionality of the inputs. For sharing, I am not sure how feasible it is. It would be quite a mess since so many files are involved and I worked locally on my computer.",
      "votes": null
    },
    {
      "id": "1913329",
      "postDate": "08/25/2022 09:26:36",
      "content": "<p>Interesting, thanks <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a>. Another quick doubt, so were you able to integrate the pretrained Electra Generator/Discriminator or you build up your own sequential model with similar training settings as of Electra.</p>",
      "rawMarkdown": "Interesting, thanks @fritzcremer. Another quick doubt, so were you able to integrate the pretrained Electra Generator/Discriminator or you build up your own sequential model with similar training settings as of Electra.",
      "votes": null
    },
    {
      "id": "1913347",
      "postDate": "08/25/2022 09:35:34",
      "content": "<p>Oh I just used tensorflow to build my own model and made several adaptations to fit the task. The model needed to predict for each position if it is a real sample, or generated by the generator. I used this approach since I felt like dealing with the vastly different and strange distributions of features might be difficult with something like an autoencoder.</p>",
      "rawMarkdown": "Oh I just used tensorflow to build my own model and made several adaptations to fit the task. The model needed to predict for each position if it is a real sample, or generated by the generator. I used this approach since I felt like dealing with the vastly different and strange distributions of features might be difficult with something like an autoencoder.",
      "votes": null
    },
    {
      "id": "1913563",
      "postDate": "08/25/2022 11:17:50",
      "content": "<p>Thanks for the write-up. Just to be clear, you train the mini-lstm independanlty right ? </p>",
      "rawMarkdown": "Thanks for the write-up. Just to be clear, you train the mini-lstm independanlty right ?",
      "votes": null
    },
    {
      "id": "1914208",
      "postDate": "08/25/2022 20:35:52",
      "content": "<p>Thank you for sharing! <br>\nI was wondering if you could give a bit more detail on the pre-training below. <br>\n<code>For the NN I have first pre-trained a transformer-model on all data which imputed random missing values.</code> <br>\nI had much difficulty figuring out a proper representation for missing values in my NN models so I am very happy to see there's a solution here. How did you build your training sample and target? And how did you structure this transformer model? </p>",
      "rawMarkdown": "Thank you for sharing! \nI was wondering if you could give a bit more detail on the pre-training below. \n`For the NN I have first pre-trained a transformer-model on all data which imputed random missing values.` \nI had much difficulty figuring out a proper representation for missing values in my NN models so I am very happy to see there's a solution here. How did you build your training sample and target? And how did you structure this transformer model?",
      "votes": null
    },
    {
      "id": "1914238",
      "postDate": "08/25/2022 21:34:50",
      "content": "<p>Yes, that is correct! They are also very small to reduce overfitting.</p>",
      "rawMarkdown": "Yes, that is correct! They are also very small to reduce overfitting.",
      "votes": null
    },
    {
      "id": "1914265",
      "postDate": "08/25/2022 23:13:04",
      "content": "<p>If I see a similar enough competition in the future, my thought would be to find multiple diverse models to run on the 13 value sequences, save each one as a feature. Something like all four of LGBM, LSTM, KNN, linear regression. </p>",
      "rawMarkdown": "If I see a similar enough competition in the future, my thought would be to find multiple diverse models to run on the 13 value sequences, save each one as a feature. Something like all four of LGBM, LSTM, KNN, linear regression.",
      "votes": null
    },
    {
      "id": "1914784",
      "postDate": "08/26/2022 12:24:58",
      "content": "<p>This is the entire model for imputing:</p>\n<p>input_x1 - numericals<br>\ninput_x1_nan - binary mask of missing values<br>\ninput_x2 - categoricals<br>\ninput_mask - mask for the sequence length (usually a value of 13)</p>\n<pre><code>def get_model(seed=0):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input((max_len, dim_f1), dtype=tf.float32)\n    input_x1_nan = Input((max_len, dim_f1), dtype=tf.float16)\n    input_x2 = Input((max_len, dim_f2), dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh0 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 0]))(input_x2[:, :,  0])\n    oh1 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 1]))(input_x2[:, :,  1])\n    oh2 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 2]))(input_x2[:, :,  2])\n    oh3 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 3]))(input_x2[:, :,  3])\n    oh4 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 4]))(input_x2[:, :,  4])\n    oh5 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 5]))(input_x2[:, :,  5])\n    oh6 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 6]))(input_x2[:, :,  6])\n    oh7 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 7]))(input_x2[:, :,  7])\n    oh8 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 8]))(input_x2[:, :,  8])\n    oh9 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 9]))(input_x2[:, :,  9])\n    ohA = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[10]))(input_x2[:, :, 10])\n\n    tile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\n    mask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n\n    dense_nan = layers.Dense(transformer_dim)(input_x1_nan)\n    dense_nan = dense_nan + positional_embedding_tf\n    nan_attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense_nan, dense_nan, attention_mask=mask)\n    nan_attention = tf.cast(nan_attention, tf.float16) * tf.cast(tf.expand_dims(input_mask, -1), tf.float16)\n    nan_attention = layers.Dense(dim_f1)(nan_attention)\n\n    oh_features = layers.Concatenate(axis=2)([oh0, oh1, oh2, oh3, oh4, oh5, oh6, oh7, oh8, oh9, ohA])\n    features = layers.Concatenate(axis=2)([tf.clip_by_value(input_x1 * tf.cast(1-input_x1_nan, tf.float32) + nan_attention, -clip_val, clip_val), oh_features])\n\n    dense = layers.Dense(1024, activation=\"elu\")(features)\n    dense = layers.Dense(transformer_dim, activation=\"elu\")(dense)\n    dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n\n    dense = dense + positional_embedding_tf\n\n    attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense, dense, attention_mask=mask)\n    attention = layers.Multiply()([attention, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n    add = layers.Add()([dense, attention])\n    norm = layers.LayerNormalization(axis=2)(add)\n    dense = layers.Dense(512, activation=\"elu\")(norm)\n    dense = layers.Dense(256, activation=\"elu\")(dense)\n    output = layers.Dense(dim_f1)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=output)\n    model_emb = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=norm)\n\n    return model, model_emb \n</code></pre>\n<p>Basically, I first convert the categoricals to one hot vectors. I set all the masked values to 0 and have an attention layers that uses the binary mask as input, and adds the results to the features (the idea being, that the model can do some pre-processing depending on what and how many values need to be imputed. Then I pass the features into a dense layer and add a positional encoding (taken from the Attention Is All You Need paper). Then I have the actual attention and predict all values and use a custom loss which only counts the masked positions. Also, I only impute the non-categorical values since I found the categorical values hard to deal with. </p>\n<p>In the final model, I cut the pre-trained model off after the norm (therefore the function returns 2 models) and add some other features.</p>",
      "rawMarkdown": "This is the entire model for imputing:\n\ninput_x1 - numericals\ninput_x1_nan - binary mask of missing values\ninput_x2 - categoricals\ninput_mask - mask for the sequence length (usually a value of 13)\n\n```\ndef get_model(seed=0):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input((max_len, dim_f1), dtype=tf.float32)\n    input_x1_nan = Input((max_len, dim_f1), dtype=tf.float16)\n    input_x2 = Input((max_len, dim_f2), dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh0 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 0]))(input_x2[:, :,  0])\n    oh1 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 1]))(input_x2[:, :,  1])\n    oh2 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 2]))(input_x2[:, :,  2])\n    oh3 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 3]))(input_x2[:, :,  3])\n    oh4 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 4]))(input_x2[:, :,  4])\n    oh5 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 5]))(input_x2[:, :,  5])\n    oh6 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 6]))(input_x2[:, :,  6])\n    oh7 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 7]))(input_x2[:, :,  7])\n    oh8 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 8]))(input_x2[:, :,  8])\n    oh9 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 9]))(input_x2[:, :,  9])\n    ohA = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[10]))(input_x2[:, :, 10])\n\n    tile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\n    mask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n\n    dense_nan = layers.Dense(transformer_dim)(input_x1_nan)\n    dense_nan = dense_nan + positional_embedding_tf\n    nan_attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense_nan, dense_nan, attention_mask=mask)\n    nan_attention = tf.cast(nan_attention, tf.float16) * tf.cast(tf.expand_dims(input_mask, -1), tf.float16)\n    nan_attention = layers.Dense(dim_f1)(nan_attention)\n\n    oh_features = layers.Concatenate(axis=2)([oh0, oh1, oh2, oh3, oh4, oh5, oh6, oh7, oh8, oh9, ohA])\n    features = layers.Concatenate(axis=2)([tf.clip_by_value(input_x1 * tf.cast(1-input_x1_nan, tf.float32) + nan_attention, -clip_val, clip_val), oh_features])\n\n    dense = layers.Dense(1024, activation=\"elu\")(features)\n    dense = layers.Dense(transformer_dim, activation=\"elu\")(dense)\n    dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n\n    dense = dense + positional_embedding_tf\n\n    attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense, dense, attention_mask=mask)\n    attention = layers.Multiply()([attention, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n    add = layers.Add()([dense, attention])\n    norm = layers.LayerNormalization(axis=2)(add)\n    dense = layers.Dense(512, activation=\"elu\")(norm)\n    dense = layers.Dense(256, activation=\"elu\")(dense)\n    output = layers.Dense(dim_f1)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=output)\n    model_emb = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=norm)\n\n    return model, model_emb \n```\n\nBasically, I first convert the categoricals to one hot vectors. I set all the masked values to 0 and have an attention layers that uses the binary mask as input, and adds the results to the features (the idea being, that the model can do some pre-processing depending on what and how many values need to be imputed. Then I pass the features into a dense layer and add a positional encoding (taken from the Attention Is All You Need paper). Then I have the actual attention and predict all values and use a custom loss which only counts the masked positions. Also, I only impute the non-categorical values since I found the categorical values hard to deal with. \n\nIn the final model, I cut the pre-trained model off after the norm (therefore the function returns 2 models) and add some other features.",
      "votes": null
    },
    {
      "id": "1914804",
      "postDate": "08/26/2022 12:42:44",
      "content": "<p><a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> thank you so much for sharing the code! I will study it in more detail! </p>",
      "rawMarkdown": "fritzcremer thank you so much for sharing the code! I will study it in more detail!",
      "votes": null
    },
    {
      "id": "1914833",
      "postDate": "08/26/2022 13:18:36",
      "content": "<p>Sure, no problem :) Some things probably look a bit confusing.<br>\nFor example, lines like these<br>\n<code>dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])</code><br>\nare there to pad the sequence with 0s when there are less then 13 statements.<br>\nAnd these 2 lines</p>\n<pre><code>tile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\nmask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n</code></pre>\n<p>are there to create the attention mask (again, for dealing with shorter sequences).</p>",
      "rawMarkdown": "Sure, no problem :) Some things probably look a bit confusing.\nFor example, lines like these\n`dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])`\nare there to pad the sequence with 0s when there are less then 13 statements.\nAnd these 2 lines\n```\ntile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\nmask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n```\nare there to create the attention mask (again, for dealing with shorter sequences).",
      "votes": null
    },
    {
      "id": "1918446",
      "postDate": "08/29/2022 15:23:02",
      "content": "<p>Hey, check it, you have a gold medal now. congratulations!</p>",
      "rawMarkdown": "Hey, check it, you have a gold medal now. congratulations!",
      "votes": null
    },
    {
      "id": "1918495",
      "postDate": "08/29/2022 16:03:22",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> achieving solo gold medal. Well done!</p>\n<p>I like your mini-lstm idea, that's a great idea. Did you train the mini-lstms using both train and test data? When predicting 1 feature, did you use all other features as input, or just the one feature as input?</p>",
      "rawMarkdown": "Congratulations @fritzcremer achieving solo gold medal. Well done!\n\nI like your mini-lstm idea, that's a great idea. Did you train the mini-lstms using both train and test data? When predicting 1 feature, did you use all other features as input, or just the one feature as input?",
      "votes": null
    },
    {
      "id": "1918575",
      "postDate": "08/29/2022 17:14:40",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! I just noticed by reading your post that I slipped into Gold :D I hope it stays that way (not sure what happens if one of the leaderboard redactions will be reversed?).</p>\n<p>About the mini-lstm, the Idea was that I still predict the original target, but the model only uses one feature as input each time. But the other Idea might also be interesting since it allows training on all the data.</p>\n<p>This is the code for the model that handles single numerical features:</p>\n<pre><code>def get_small_model(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(max_len, dtype=tf.float32) \n    input_x1_nans = Input(max_len, dtype=tf.float32)\n    input_mask = Input(max_len)\n\n    lag = input_x1[:, 1:] - input_x1[:, :-1]\n    lag = tf.concat([lag[:, :1], lag], axis=1) * input_mask # Pad the lag features to length 13\n\n    fmin = tf.reduce_min(input_x1 + 1000 * (1 - input_mask), axis=1, keepdims=True) # Min of the feature (+1000 on the padded part to ignore it)\n    fmax = tf.reduce_max(input_x1 - 1000 * (1 - input_mask), axis=1, keepdims=True) # Max of the feature (-1000 on the padded part to ignore it)\n    fstd = tf.sqrt(tf.math.reduce_sum(tf.square(input_x1), axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True)) # Std of the feature\n    fmean = tf.math.reduce_sum(input_x1, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Mean of the feature\n    fmean_nans = tf.math.reduce_sum(input_x1_nans, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Fraction of nan values\n\n    dense = layers.Dense(16, activation=\"swish\")(tf.stack([input_x1, lag, input_mask], axis=2)) # Pass features through dense\n\n    lstm = layers.LSTM(10)(dense)\n\n    concat = layers.Concatenate()([lag[:, -1:], input_x1[:, -1:], fmin, fmax, fstd, fmean, fmean_nans, lstm])\n    dense = layers.Dense(8, activation=\"swish\")(concat)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nans, input_mask], outputs=output)\n\n    return model\n</code></pre>\n<p>and the model that handles categorical features:</p>\n<pre><code>def get_small_model_cat(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(13, dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh = layers.Lambda(lambda x : tf.one_hot(x, depth=20))(input_x1)\n\n    lstm = layers.LSTM(10)(oh)\n\n    dense = layers.Dense(8, activation=\"swish\")(lstm)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_mask], outputs=output)\n\n    return model\n</code></pre>\n<p>Since the 179 training-runs (for each feature) took quite a long time, I only trained with one validation fold as hold-out set and did some precautions to not overfit (very small model, max 16 epochs). This way I was also able to get quite a nice overview of the single feature predictive power:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6735218%2F9fa502558e06c614ae2135703b47612b%2Fsingle%20feature%20predicitveness.png?generation=1661792977828903&amp;alt=media\" alt=\"\"></p>\n<p>Each dot is a feature and the colors correspond to the different features groups:</p>\n<p>S - red<br>\nP - green<br>\nD - blue<br>\nB - purple<br>\nR - black</p>\n<p>Further optimizing the mini-lstms might help but along with my pre-trained impute model, I think they gave me the largest boost.</p>",
      "rawMarkdown": "Thank you @cdeotte! I just noticed by reading your post that I slipped into Gold :D I hope it stays that way (not sure what happens if one of the leaderboard redactions will be reversed?).\n\nAbout the mini-lstm, the Idea was that I still predict the original target, but the model only uses one feature as input each time. But the other Idea might also be interesting since it allows training on all the data.\n\nThis is the code for the model that handles single numerical features:\n\n```\ndef get_small_model(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(max_len, dtype=tf.float32) \n    input_x1_nans = Input(max_len, dtype=tf.float32)\n    input_mask = Input(max_len)\n\n    lag = input_x1[:, 1:] - input_x1[:, :-1]\n    lag = tf.concat([lag[:, :1], lag], axis=1) * input_mask # Pad the lag features to length 13\n\n    fmin = tf.reduce_min(input_x1 + 1000 * (1 - input_mask), axis=1, keepdims=True) # Min of the feature (+1000 on the padded part to ignore it)\n    fmax = tf.reduce_max(input_x1 - 1000 * (1 - input_mask), axis=1, keepdims=True) # Max of the feature (-1000 on the padded part to ignore it)\n    fstd = tf.sqrt(tf.math.reduce_sum(tf.square(input_x1), axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True)) # Std of the feature\n    fmean = tf.math.reduce_sum(input_x1, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Mean of the feature\n    fmean_nans = tf.math.reduce_sum(input_x1_nans, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Fraction of nan values\n\n    dense = layers.Dense(16, activation=\"swish\")(tf.stack([input_x1, lag, input_mask], axis=2)) # Pass features through dense\n\n    lstm = layers.LSTM(10)(dense)\n\n    concat = layers.Concatenate()([lag[:, -1:], input_x1[:, -1:], fmin, fmax, fstd, fmean, fmean_nans, lstm])\n    dense = layers.Dense(8, activation=\"swish\")(concat)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nans, input_mask], outputs=output)\n\n    return model\n```\n\nand the model that handles categorical features:\n\n```\ndef get_small_model_cat(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(13, dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh = layers.Lambda(lambda x : tf.one_hot(x, depth=20))(input_x1)\n\n    lstm = layers.LSTM(10)(oh)\n\n    dense = layers.Dense(8, activation=\"swish\")(lstm)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_mask], outputs=output)\n\n    return model\n```\n\nSince the 179 training-runs (for each feature) took quite a long time, I only trained with one validation fold as hold-out set and did some precautions to not overfit (very small model, max 16 epochs). This way I was also able to get quite a nice overview of the single feature predictive power:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6735218%2F9fa502558e06c614ae2135703b47612b%2Fsingle%20feature%20predicitveness.png?generation=1661792977828903&alt=media)\n\nEach dot is a feature and the colors correspond to the different features groups:\n\nS - red\nP - green\nD - blue\nB - purple\nR - black\n\nFurther optimizing the mini-lstms might help but along with my pre-trained impute model, I think they gave me the largest boost.",
      "votes": null
    },
    {
      "id": "1920885",
      "postDate": "08/31/2022 13:06:21",
      "content": "<p>Congratulations! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Congratulations! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    },
    {
      "id": "1924091",
      "postDate": "09/02/2022 17:53:10",
      "content": "<p>Yes, thanks :D I couldn't believe it ^^</p>",
      "rawMarkdown": "Yes, thanks :D I couldn't believe it ^^",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1912849,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "08/25/2022 02:08:38",
      "content": "<p>Great work! Congrats on your achievement!</p>\n<p>We used concept which is pretty similar to your mini-LSTM, we found that train lightgbm on each feature produce a better results.</p>\n<p>You can find the notebook <a href=\"https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431\" target=\"_blank\">here</a><br>\nAny chance that you will share your solution as a notebook?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913229,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/25/2022 08:41:08",
          "content": "<p>Oh that's interesting, I though about an LGBM approach as well but thought LSTM might outperform them, because of the sequence structure and relatively low dimensionality of the inputs. For sharing, I am not sure how feasible it is. It would be quite a mess since so many files are involved and I worked locally on my computer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914265,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/25/2022 23:13:04",
          "content": "<p>If I see a similar enough competition in the future, my thought would be to find multiple diverse models to run on the 13 value sequences, save each one as a feature. Something like all four of LGBM, LSTM, KNN, linear regression. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1912934,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "08/25/2022 03:55:44",
      "content": "<p>Congratulations and thanks for sharing the solution. Can you please elaborate on the Electra Pretraining part?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1913175,
          "author_name": "nyleve",
          "author_url": "",
          "post_date": "08/25/2022 07:44:32",
          "content": "<p>I second the request </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1913227,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/25/2022 08:38:45",
          "content": "<p>Yes, so I though about a way to generate embeddings by pre-training on all data. I was inspired by the training process of Google's language model ELECTRA. I used a generator and a discriminator. The generator learned to produce sequence elements (e.g. the entire last statement) and swapped them with the actual data. The discriminator needed to decide, which of the 13 sequence elements have been created by the generator (e.g. 5, 7, 10, 11 have been produced by the generator). This produced embeddings in the last hidden layer that had quite a decent predictive power (though way too weak to use on it's own), so I experimented with merging those features with the NN and the LGBM. With the NN I saw a minor improvement and with the LGBM no improvement (this was probably due to noise since each training went a bit different anyways). For the NN it also could have been noise but I just kept it in.</p>\n<p>And actually I did not add the features itself, but rather I trained another NN to produce a single probabiltiy from those embeddings, and fed this single feature to the NN. Overall, my pipeline feels a bit complicated and things like these could probably be removed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1913329,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "08/25/2022 09:26:36",
          "content": "<p>Interesting, thanks <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a>. Another quick doubt, so were you able to integrate the pretrained Electra Generator/Discriminator or you build up your own sequential model with similar training settings as of Electra.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1913347,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/25/2022 09:35:34",
          "content": "<p>Oh I just used tensorflow to build my own model and made several adaptations to fit the task. The model needed to predict for each position if it is a real sample, or generated by the generator. I used this approach since I felt like dealing with the vastly different and strange distributions of features might be difficult with something like an autoencoder.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1913563,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "08/25/2022 11:17:50",
      "content": "<p>Thanks for the write-up. Just to be clear, you train the mini-lstm independanlty right ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1914238,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/25/2022 21:34:50",
          "content": "<p>Yes, that is correct! They are also very small to reduce overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1914208,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "08/25/2022 20:35:52",
      "content": "<p>Thank you for sharing! <br>\nI was wondering if you could give a bit more detail on the pre-training below. <br>\n<code>For the NN I have first pre-trained a transformer-model on all data which imputed random missing values.</code> <br>\nI had much difficulty figuring out a proper representation for missing values in my NN models so I am very happy to see there's a solution here. How did you build your training sample and target? And how did you structure this transformer model? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1914784,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/26/2022 12:24:58",
          "content": "<p>This is the entire model for imputing:</p>\n<p>input_x1 - numericals<br>\ninput_x1_nan - binary mask of missing values<br>\ninput_x2 - categoricals<br>\ninput_mask - mask for the sequence length (usually a value of 13)</p>\n<pre><code>def get_model(seed=0):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input((max_len, dim_f1), dtype=tf.float32)\n    input_x1_nan = Input((max_len, dim_f1), dtype=tf.float16)\n    input_x2 = Input((max_len, dim_f2), dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh0 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 0]))(input_x2[:, :,  0])\n    oh1 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 1]))(input_x2[:, :,  1])\n    oh2 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 2]))(input_x2[:, :,  2])\n    oh3 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 3]))(input_x2[:, :,  3])\n    oh4 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 4]))(input_x2[:, :,  4])\n    oh5 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 5]))(input_x2[:, :,  5])\n    oh6 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 6]))(input_x2[:, :,  6])\n    oh7 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 7]))(input_x2[:, :,  7])\n    oh8 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 8]))(input_x2[:, :,  8])\n    oh9 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 9]))(input_x2[:, :,  9])\n    ohA = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[10]))(input_x2[:, :, 10])\n\n    tile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\n    mask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n\n    dense_nan = layers.Dense(transformer_dim)(input_x1_nan)\n    dense_nan = dense_nan + positional_embedding_tf\n    nan_attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense_nan, dense_nan, attention_mask=mask)\n    nan_attention = tf.cast(nan_attention, tf.float16) * tf.cast(tf.expand_dims(input_mask, -1), tf.float16)\n    nan_attention = layers.Dense(dim_f1)(nan_attention)\n\n    oh_features = layers.Concatenate(axis=2)([oh0, oh1, oh2, oh3, oh4, oh5, oh6, oh7, oh8, oh9, ohA])\n    features = layers.Concatenate(axis=2)([tf.clip_by_value(input_x1 * tf.cast(1-input_x1_nan, tf.float32) + nan_attention, -clip_val, clip_val), oh_features])\n\n    dense = layers.Dense(1024, activation=\"elu\")(features)\n    dense = layers.Dense(transformer_dim, activation=\"elu\")(dense)\n    dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n\n    dense = dense + positional_embedding_tf\n\n    attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense, dense, attention_mask=mask)\n    attention = layers.Multiply()([attention, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n    add = layers.Add()([dense, attention])\n    norm = layers.LayerNormalization(axis=2)(add)\n    dense = layers.Dense(512, activation=\"elu\")(norm)\n    dense = layers.Dense(256, activation=\"elu\")(dense)\n    output = layers.Dense(dim_f1)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=output)\n    model_emb = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=norm)\n\n    return model, model_emb \n</code></pre>\n<p>Basically, I first convert the categoricals to one hot vectors. I set all the masked values to 0 and have an attention layers that uses the binary mask as input, and adds the results to the features (the idea being, that the model can do some pre-processing depending on what and how many values need to be imputed. Then I pass the features into a dense layer and add a positional encoding (taken from the Attention Is All You Need paper). Then I have the actual attention and predict all values and use a custom loss which only counts the masked positions. Also, I only impute the non-categorical values since I found the categorical values hard to deal with. </p>\n<p>In the final model, I cut the pre-trained model off after the norm (therefore the function returns 2 models) and add some other features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914804,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "08/26/2022 12:42:44",
          "content": "<p><a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> thank you so much for sharing the code! I will study it in more detail! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914833,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/26/2022 13:18:36",
          "content": "<p>Sure, no problem :) Some things probably look a bit confusing.<br>\nFor example, lines like these<br>\n<code>dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])</code><br>\nare there to pad the sequence with 0s when there are less then 13 statements.<br>\nAnd these 2 lines</p>\n<pre><code>tile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\nmask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n</code></pre>\n<p>are there to create the attention mask (again, for dealing with shorter sequences).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1918446,
      "author_name": "meli19",
      "author_url": "",
      "post_date": "08/29/2022 15:23:02",
      "content": "<p>Hey, check it, you have a gold medal now. congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1924091,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "09/02/2022 17:53:10",
          "content": "<p>Yes, thanks :D I couldn't believe it ^^</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1918495,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/29/2022 16:03:22",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/fritzcremer\" target=\"_blank\">@fritzcremer</a> achieving solo gold medal. Well done!</p>\n<p>I like your mini-lstm idea, that's a great idea. Did you train the mini-lstms using both train and test data? When predicting 1 feature, did you use all other features as input, or just the one feature as input?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1918575,
          "author_name": "fritzcremer",
          "author_url": "",
          "post_date": "08/29/2022 17:14:40",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! I just noticed by reading your post that I slipped into Gold :D I hope it stays that way (not sure what happens if one of the leaderboard redactions will be reversed?).</p>\n<p>About the mini-lstm, the Idea was that I still predict the original target, but the model only uses one feature as input each time. But the other Idea might also be interesting since it allows training on all the data.</p>\n<p>This is the code for the model that handles single numerical features:</p>\n<pre><code>def get_small_model(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(max_len, dtype=tf.float32) \n    input_x1_nans = Input(max_len, dtype=tf.float32)\n    input_mask = Input(max_len)\n\n    lag = input_x1[:, 1:] - input_x1[:, :-1]\n    lag = tf.concat([lag[:, :1], lag], axis=1) * input_mask # Pad the lag features to length 13\n\n    fmin = tf.reduce_min(input_x1 + 1000 * (1 - input_mask), axis=1, keepdims=True) # Min of the feature (+1000 on the padded part to ignore it)\n    fmax = tf.reduce_max(input_x1 - 1000 * (1 - input_mask), axis=1, keepdims=True) # Max of the feature (-1000 on the padded part to ignore it)\n    fstd = tf.sqrt(tf.math.reduce_sum(tf.square(input_x1), axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True)) # Std of the feature\n    fmean = tf.math.reduce_sum(input_x1, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Mean of the feature\n    fmean_nans = tf.math.reduce_sum(input_x1_nans, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Fraction of nan values\n\n    dense = layers.Dense(16, activation=\"swish\")(tf.stack([input_x1, lag, input_mask], axis=2)) # Pass features through dense\n\n    lstm = layers.LSTM(10)(dense)\n\n    concat = layers.Concatenate()([lag[:, -1:], input_x1[:, -1:], fmin, fmax, fstd, fmean, fmean_nans, lstm])\n    dense = layers.Dense(8, activation=\"swish\")(concat)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nans, input_mask], outputs=output)\n\n    return model\n</code></pre>\n<p>and the model that handles categorical features:</p>\n<pre><code>def get_small_model_cat(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(13, dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh = layers.Lambda(lambda x : tf.one_hot(x, depth=20))(input_x1)\n\n    lstm = layers.LSTM(10)(oh)\n\n    dense = layers.Dense(8, activation=\"swish\")(lstm)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_mask], outputs=output)\n\n    return model\n</code></pre>\n<p>Since the 179 training-runs (for each feature) took quite a long time, I only trained with one validation fold as hold-out set and did some precautions to not overfit (very small model, max 16 epochs). This way I was also able to get quite a nice overview of the single feature predictive power:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6735218%2F9fa502558e06c614ae2135703b47612b%2Fsingle%20feature%20predicitveness.png?generation=1661792977828903&amp;alt=media\" alt=\"\"></p>\n<p>Each dot is a feature and the colors correspond to the different features groups:</p>\n<p>S - red<br>\nP - green<br>\nD - blue<br>\nB - purple<br>\nR - black</p>\n<p>Further optimizing the mini-lstms might help but along with my pre-trained impute model, I think they gave me the largest boost.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1920885,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "08/31/2022 13:06:21",
      "content": "<p>Congratulations! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912797": "Thanks to AMEX for hosting this competition!\nI just want to give a quick overview of the stuff I did in this competition.\n\n# Main concepts\n\nI used an NN based models and LGBM based models for this competition. \n#### Pre-processing \nBecause of some features with very spread out distributions, I figured that log-transforming several features for the NN models might make sense. I log-transformed a good fraction of the features. I decided to transform or not by gut feeling after looking at the histogram of each feature.\n\nI noticed 2 features that behaved weirdly in private data, so I removed them for all NN models.\nThose features **D_59**, **D_86**. **D_59** even had a different private test data distribution than public test data distribution.\n\n#### Mini-LSTMs\nFor both model types, I constructed new features by training a simple LSTM on every single feature (13x1 input dimension). Those models still predict the original target.\nThis results in 189 different models which generate probabilities given the sequence of only a single feature.\n\n#### LGBMs\nFor the LGBM based model, I used the great [notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977) by @ragnar123 as a starting point. I added the previously generated features from the LSTM models and changed tweaked a few hyperparameters. I also did one training with all data.\n\n#### NN\nFor the NN I have first pre-trained a transformer-model on all data which imputed random missing values. I then embedded this pre-trained model into an NN which also used several other features (e.g. last sequence element) and trained it. This improved the NN performance quite a bit. I also passed the Mini-LSTM model features into the NN.\n\n#### Pseudo-labeling\nI also trained the previous NN on pseudo labeled test data (labeled by all other previous models), which gave this nn a score of 0.799 on public LB. This was however difficult the validate, since the pseudo labels created a heavy leak. Therefore I validated this model only on the public LB data. This helped to further increase my score.\n\n#### Final model\nIn the end I ensembled the NNs and LGBMs with roughly 30/70 (NN/LGBM) weight.\nI also included this [public submission](https://www.kaggle.com/code/swimmy/tuffline-amex-anotherfeaturelgbm) in the ensemble.\n\n# Things that did not work\nOr at least things that didn't help much.\n\nI tried using pre-trained autoencoders, ELECTRA style pretraining, Tabnet, CNN, LSTM, XGB, other losses (pearson, mse), and some smaller stuff...\nI tried **a lot** of things with some of these but it just didn't really help. I actually used embeddings generated by the ELECTRA style pre-trained transformer model for my NN, but I think this only contributes to a minor improvement.\n\nAnyway, that's basically it. I am sure all of it 100-200 hours and I feel like there is still a lot of room for improvement (e.g. I did not test enough features for LGBMs, could still train more models since I only use like 25 single models in the final version). I focused a lot on playing with NN in this competition, but a lot of stuff just wouldn't work.",
    "1912849": "Great work! Congrats on your achievement!\n\nWe used concept which is pretty similar to your mini-LSTM, we found that train lightgbm on each feature produce a better results.\n\nYou can find the notebook [here](https://www.kaggle.com/code/pavelvod/27-place-sequentialencoder?scriptVersionId=104154431)\nAny chance that you will share your solution as a notebook?",
    "1912934": "Congratulations and thanks for sharing the solution. Can you please elaborate on the Electra Pretraining part?",
    "1913175": "I second the request",
    "1913227": "Yes, so I though about a way to generate embeddings by pre-training on all data. I was inspired by the training process of Google's language model ELECTRA. I used a generator and a discriminator. The generator learned to produce sequence elements (e.g. the entire last statement) and swapped them with the actual data. The discriminator needed to decide, which of the 13 sequence elements have been created by the generator (e.g. 5, 7, 10, 11 have been produced by the generator). This produced embeddings in the last hidden layer that had quite a decent predictive power (though way too weak to use on it's own), so I experimented with merging those features with the NN and the LGBM. With the NN I saw a minor improvement and with the LGBM no improvement (this was probably due to noise since each training went a bit different anyways). For the NN it also could have been noise but I just kept it in.\n\nAnd actually I did not add the features itself, but rather I trained another NN to produce a single probabiltiy from those embeddings, and fed this single feature to the NN. Overall, my pipeline feels a bit complicated and things like these could probably be removed.",
    "1913229": "Oh that's interesting, I though about an LGBM approach as well but thought LSTM might outperform them, because of the sequence structure and relatively low dimensionality of the inputs. For sharing, I am not sure how feasible it is. It would be quite a mess since so many files are involved and I worked locally on my computer.",
    "1913329": "Interesting, thanks @fritzcremer. Another quick doubt, so were you able to integrate the pretrained Electra Generator/Discriminator or you build up your own sequential model with similar training settings as of Electra.",
    "1913347": "Oh I just used tensorflow to build my own model and made several adaptations to fit the task. The model needed to predict for each position if it is a real sample, or generated by the generator. I used this approach since I felt like dealing with the vastly different and strange distributions of features might be difficult with something like an autoencoder.",
    "1913563": "Thanks for the write-up. Just to be clear, you train the mini-lstm independanlty right ?",
    "1914208": "Thank you for sharing! \nI was wondering if you could give a bit more detail on the pre-training below. \n`For the NN I have first pre-trained a transformer-model on all data which imputed random missing values.` \nI had much difficulty figuring out a proper representation for missing values in my NN models so I am very happy to see there's a solution here. How did you build your training sample and target? And how did you structure this transformer model?",
    "1914238": "Yes, that is correct! They are also very small to reduce overfitting.",
    "1914265": "If I see a similar enough competition in the future, my thought would be to find multiple diverse models to run on the 13 value sequences, save each one as a feature. Something like all four of LGBM, LSTM, KNN, linear regression.",
    "1914784": "This is the entire model for imputing:\n\ninput_x1 - numericals\ninput_x1_nan - binary mask of missing values\ninput_x2 - categoricals\ninput_mask - mask for the sequence length (usually a value of 13)\n\n```\ndef get_model(seed=0):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input((max_len, dim_f1), dtype=tf.float32)\n    input_x1_nan = Input((max_len, dim_f1), dtype=tf.float16)\n    input_x2 = Input((max_len, dim_f2), dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh0 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 0]))(input_x2[:, :,  0])\n    oh1 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 1]))(input_x2[:, :,  1])\n    oh2 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 2]))(input_x2[:, :,  2])\n    oh3 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 3]))(input_x2[:, :,  3])\n    oh4 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 4]))(input_x2[:, :,  4])\n    oh5 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 5]))(input_x2[:, :,  5])\n    oh6 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 6]))(input_x2[:, :,  6])\n    oh7 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 7]))(input_x2[:, :,  7])\n    oh8 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 8]))(input_x2[:, :,  8])\n    oh9 = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[ 9]))(input_x2[:, :,  9])\n    ohA = layers.Lambda(lambda x : tf.one_hot(x, depth=cat_counts[10]))(input_x2[:, :, 10])\n\n    tile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\n    mask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n\n    dense_nan = layers.Dense(transformer_dim)(input_x1_nan)\n    dense_nan = dense_nan + positional_embedding_tf\n    nan_attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense_nan, dense_nan, attention_mask=mask)\n    nan_attention = tf.cast(nan_attention, tf.float16) * tf.cast(tf.expand_dims(input_mask, -1), tf.float16)\n    nan_attention = layers.Dense(dim_f1)(nan_attention)\n\n    oh_features = layers.Concatenate(axis=2)([oh0, oh1, oh2, oh3, oh4, oh5, oh6, oh7, oh8, oh9, ohA])\n    features = layers.Concatenate(axis=2)([tf.clip_by_value(input_x1 * tf.cast(1-input_x1_nan, tf.float32) + nan_attention, -clip_val, clip_val), oh_features])\n\n    dense = layers.Dense(1024, activation=\"elu\")(features)\n    dense = layers.Dense(transformer_dim, activation=\"elu\")(dense)\n    dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n\n    dense = dense + positional_embedding_tf\n\n    attention = layers.MultiHeadAttention(num_heads=4, key_dim=32, value_dim=32, output_shape=transformer_dim)(dense, dense, attention_mask=mask)\n    attention = layers.Multiply()([attention, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])\n    add = layers.Add()([dense, attention])\n    norm = layers.LayerNormalization(axis=2)(add)\n    dense = layers.Dense(512, activation=\"elu\")(norm)\n    dense = layers.Dense(256, activation=\"elu\")(dense)\n    output = layers.Dense(dim_f1)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=output)\n    model_emb = Model(inputs=[input_x1, input_x1_nan, input_x2, input_mask], outputs=norm)\n\n    return model, model_emb \n```\n\nBasically, I first convert the categoricals to one hot vectors. I set all the masked values to 0 and have an attention layers that uses the binary mask as input, and adds the results to the features (the idea being, that the model can do some pre-processing depending on what and how many values need to be imputed. Then I pass the features into a dense layer and add a positional encoding (taken from the Attention Is All You Need paper). Then I have the actual attention and predict all values and use a custom loss which only counts the masked positions. Also, I only impute the non-categorical values since I found the categorical values hard to deal with. \n\nIn the final model, I cut the pre-trained model off after the norm (therefore the function returns 2 models) and add some other features.",
    "1914804": "fritzcremer thank you so much for sharing the code! I will study it in more detail!",
    "1914833": "Sure, no problem :) Some things probably look a bit confusing.\nFor example, lines like these\n`dense = layers.Multiply()([dense, tf.tile(tf.expand_dims(input_mask, -1), (1, 1, transformer_dim))])`\nare there to pad the sequence with 0s when there are less then 13 statements.\nAnd these 2 lines\n```\ntile = layers.Lambda(lambda x: tf.tile(tf.expand_dims(x, -1), (1, 1, 13)))\nmask = layers.Multiply()([tile(input_mask), tf.transpose(tile(input_mask), perm=(0, 2, 1))])\n```\nare there to create the attention mask (again, for dealing with shorter sequences).",
    "1918446": "Hey, check it, you have a gold medal now. congratulations!",
    "1918495": "Congratulations @fritzcremer achieving solo gold medal. Well done!\n\nI like your mini-lstm idea, that's a great idea. Did you train the mini-lstms using both train and test data? When predicting 1 feature, did you use all other features as input, or just the one feature as input?",
    "1918575": "Thank you @cdeotte! I just noticed by reading your post that I slipped into Gold :D I hope it stays that way (not sure what happens if one of the leaderboard redactions will be reversed?).\n\nAbout the mini-lstm, the Idea was that I still predict the original target, but the model only uses one feature as input each time. But the other Idea might also be interesting since it allows training on all the data.\n\nThis is the code for the model that handles single numerical features:\n\n```\ndef get_small_model(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(max_len, dtype=tf.float32) \n    input_x1_nans = Input(max_len, dtype=tf.float32)\n    input_mask = Input(max_len)\n\n    lag = input_x1[:, 1:] - input_x1[:, :-1]\n    lag = tf.concat([lag[:, :1], lag], axis=1) * input_mask # Pad the lag features to length 13\n\n    fmin = tf.reduce_min(input_x1 + 1000 * (1 - input_mask), axis=1, keepdims=True) # Min of the feature (+1000 on the padded part to ignore it)\n    fmax = tf.reduce_max(input_x1 - 1000 * (1 - input_mask), axis=1, keepdims=True) # Max of the feature (-1000 on the padded part to ignore it)\n    fstd = tf.sqrt(tf.math.reduce_sum(tf.square(input_x1), axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True)) # Std of the feature\n    fmean = tf.math.reduce_sum(input_x1, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Mean of the feature\n    fmean_nans = tf.math.reduce_sum(input_x1_nans, axis=1, keepdims=True) / tf.math.reduce_sum(input_mask, axis=1, keepdims=True) # Fraction of nan values\n\n    dense = layers.Dense(16, activation=\"swish\")(tf.stack([input_x1, lag, input_mask], axis=2)) # Pass features through dense\n\n    lstm = layers.LSTM(10)(dense)\n\n    concat = layers.Concatenate()([lag[:, -1:], input_x1[:, -1:], fmin, fmax, fstd, fmean, fmean_nans, lstm])\n    dense = layers.Dense(8, activation=\"swish\")(concat)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_x1_nans, input_mask], outputs=output)\n\n    return model\n```\n\nand the model that handles categorical features:\n\n```\ndef get_small_model_cat(use_activation, seed):\n    tf.random.set_seed(seed)\n\n    input_x1 = Input(13, dtype=tf.int32)\n    input_mask = Input(max_len)\n\n    oh = layers.Lambda(lambda x : tf.one_hot(x, depth=20))(input_x1)\n\n    lstm = layers.LSTM(10)(oh)\n\n    dense = layers.Dense(8, activation=\"swish\")(lstm)\n    output = layers.Dense(1, activation=use_activation)(dense)\n\n    model = Model(inputs=[input_x1, input_mask], outputs=output)\n\n    return model\n```\n\nSince the 179 training-runs (for each feature) took quite a long time, I only trained with one validation fold as hold-out set and did some precautions to not overfit (very small model, max 16 epochs). This way I was also able to get quite a nice overview of the single feature predictive power:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6735218%2F9fa502558e06c614ae2135703b47612b%2Fsingle%20feature%20predicitveness.png?generation=1661792977828903&alt=media)\n\nEach dot is a feature and the colors correspond to the different features groups:\n\nS - red\nP - green\nD - blue\nB - purple\nR - black\n\nFurther optimizing the mini-lstms might help but along with my pre-trained impute model, I think they gave me the largest boost.",
    "1920885": "Congratulations! May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
    "1924091": "Yes, thanks :D I couldn't believe it ^^"
  },
  "source": "meta"
}