{
  "id": 416026,
  "title": "1st place solution: transformer and acceleration data",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416026",
  "author_name": "Urazalinov Baurzhan",
  "post_date": "2023-06-09T07:48:46.850000",
  "votes": 117,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Greetings to the Kaggle Community. In this message I want to tell you about my solution.</p>\n<p>Thanks to Kaggle for providing free GPU and TPU resources to everyone. On my graphics card (1050 Ti) I would not have achieved those results.<br>\nThanks to Google for the excellent tensorflow library.<br>\nAll of my work was done in Kaggle Notebooks and relies on TensorFlow capabilities.</p>\n<p>The key decisions that, in my opinion, led to a good result:</p>\n<ol>\n<li>Use a combination of transformer encoder and two BidirectionalLSTM layers.</li>\n<li>Use patches like VisualTransformer.</li>\n<li>Reduce the resolution of targets.</li>\n</ol>\n<p><em>How does it work?</em></p>\n<p>Suppose we have a tdcsfog sensor data series with AccV, AccML, AccAP columns and len of 5000.</p>\n<p>First, apply mean-std normalization to AccV, AccML, AccAP columns.</p>\n<pre><code> ():\n    mean = tf.math.reduce_mean(sample)\n    std = tf.math.reduce_std(sample)\n    sample = tf.math.divide_no_nan(sample-mean, std)\n\n     sample.numpy()\n</code></pre>\n<p>Then the series is zero-padded so that the final length is divisible by block_size = 15552  (or 12096 for defog). Now the series shape is (15552,  3). </p>\n<p>And create patches with the patch_size = 18 (or 14 for defog):</p>\n<pre><code>series \nseries = tf.reshape(series, shape=(CFG[] // CFG[], CFG[], )) \nseries = tf.reshape(series, shape=(CFG[] // CFG[], CFG[]*))  \n</code></pre>\n<p>Now the series shape is (864,  54). It's a model input.</p>\n<p>What to do with the StartHesitation, Turn, Walking data? Same, but apply tf.reduce_max at the end.</p>\n<pre><code>series_targets \nseries_targets = tf.reshape(series_targets, shape=(CFG[] // CFG[], CFG[], )) \nseries_targets = tf.transpose(series_targets, perm=[, , ]) \nseries_targets = tf.reduce_max(series_targets, axis=-) \n</code></pre>\n<p>Now the series shape is (864, 3). It's a model output.</p>\n<p>At the end, simply return the true resolution with tf.tile</p>\n<pre><code>predictions = model.predict(...) \npredictions = tf.expand_dims(predictions, axis=-) \npredictions = tf.transpose(predictions, perm=[, , , ]) \npredictions = tf.tile(predictions, multiples=[, , CFG[], ]) \npredictions = tf.reshape(predictions, shape=(predictions.shape[], predictions.shape[]*predictions.shape[], )) \n</code></pre>\n<h1>Details</h1>\n<p>Daily data, events.csv, subjects.csv, tasks.csv have never been used.</p>\n<p>Tdcsfog data is not used to train defog models. </p>\n<p>Defog data is not used to train tdcsfog models.</p>\n<p><em>Optimizer</em> </p>\n<pre><code>tf.keras.optimizers.Adam(learning_rate=Schedule(LEARNING_RATE, WARMUP_STEPS), beta_1=, beta_2=, epsilon=)\n</code></pre>\n<p><em>Loss function</em></p>\n<pre><code>\n\nce = tf.keras.losses.BinaryCrossentropy(reduction=)\n\n ():\n    loss = ce(tf.expand_dims(real[:, :, :], axis=-), tf.expand_dims(output, axis=-)) \n\n    mask = tf.math.multiply(real[:, :, ], real[:, :, ]) \n    mask = tf.cast(mask, dtype=loss.dtype)\n    mask = tf.expand_dims(mask, axis=-) \n    mask = tf.tile(mask, multiples=[, , ]) \n    loss *= mask \n\n     tf.reduce_sum(loss) / tf.reduce_sum(mask)\n</code></pre>\n<p><em>Model</em> </p>\n<pre><code>CFG = {: ,\n       : ,\n       : //,\n       : ,\n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\n\n\n (tf.keras.layers.Layer):\n     ():\n        ().__init__()\n\n        self.mha = tf.keras.layers.MultiHeadAttention(num_heads=CFG[], key_dim=CFG[], dropout=CFG[])\n\n        self.add = tf.keras.layers.Add()\n\n        self.layernorm = tf.keras.layers.LayerNormalization()\n\n        self.seq = tf.keras.Sequential([tf.keras.layers.Dense(CFG[], activation=),\n                                        tf.keras.layers.Dropout(CFG[]),\n                                        tf.keras.layers.Dense(CFG[]),\n                                        tf.keras.layers.Dropout(CFG[]),\n                                       ])\n\n     ():\n        attn_output = self.mha(query=x, key=x, value=x)\n        x = self.add([x, attn_output])\n        x = self.layernorm(x)\n        x = self.add([x, self.seq(x)])\n        x = self.layernorm(x)\n\n         x\n\n\n\n (tf.keras.Model):\n     ():\n        ().__init__()\n\n        self.first_linear = tf.keras.layers.Dense(CFG[])\n\n        self.add = tf.keras.layers.Add()\n\n        self.first_dropout = tf.keras.layers.Dropout(CFG[])\n\n        self.enc_layers = [EncoderLayer()  _  (CFG[])]\n\n        self.lstm_layers = [tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(CFG[], return_sequences=))  _  (CFG[])]\n\n        self.sequence_len = CFG[] // CFG[]\n        self.pos_encoding = tf.Variable(initial_value=tf.random.normal(shape=(, self.sequence_len, CFG[]), stddev=), trainable=)\n\n     (): \n        x = x /  \n        x = self.first_linear(x) \n\n         training: \n            random_pos_encoding = tf.roll(tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, , ]),\n                                          shift=tf.random.uniform(shape=(GPU_BATCH_SIZE,), minval=-self.sequence_len, maxval=, dtype=tf.int32),\n                                          axis=GPU_BATCH_SIZE * [],\n                                          )\n            x = self.add([x, random_pos_encoding])\n\n        : \n            x = self.add([x, tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, , ])])\n\n        x = self.first_dropout(x)\n\n         i  (CFG[]): x = self.enc_layers[i](x) \n         i  (CFG[]): x = self.lstm_layers[i](x) \n\n         x\n\n (tf.keras.Model):\n     ():\n        ().__init__()\n\n        self.encoder = FOGEncoder()\n        self.last_linear = tf.keras.layers.Dense()\n\n     (): \n        x = self.encoder(x) \n        x = self.last_linear(x) \n        x = tf.nn.sigmoid(x) \n\n         x\n</code></pre>\n<h1>Submission (Private Score 0.514, Public Score 0.527) consists of 8 models:</h1>\n<h3>Model 1 (tdcsfog model)</h3>\n<pre><code>CFG = {: , \n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE=\n</code></pre>\n<p>Validation subjects <br>\n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']</p>\n<p>Train 15 minutes on TPU. Validation scores:<br>\nStartHesitation AP - 0.462 Turn AP - 0.896 Walking AP - 0.470 mAP - 0.609</p>\n<h3>Model 2 (tdcsfog model)</h3>\n<pre><code>CFG = {: , \n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']</p>\n<p>Train 40 minutes on GPU. Validation scores:<br>\nStartHesitation AP - 0.481 Turn AP - 0.886 Walking AP - 0.437 mAP - 0.601</p>\n<h3>Model 3 (tdcsfog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['e39bc5', '516a67', 'af82b2', '4dc2f8', '743f4e', 'fa8764', 'a03db7', '51574c', '2d57c2']</p>\n<p>Train 11 minutes on TPU. Validation scores:<br>\nStartHesitation AP - 0.601 Turn AP - 0.857 Walking AP - 0.289 mAP - 0.582</p>\n<h3>Model 4 (tdcsfog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['5c0b8a', 'a03db7', '7fcee9', '2c98f7', '2a39f8', '4f13b4', 'af82b2', 'f686f0', '93f49f', '194d1d', '02bc69', '082f01']</p>\n<p>Train 13 minutes on TPU. Validation scores:<br>\nStartHesitation AP - 0.367 Turn AP - 0.879 Walking AP - 0.194 mAP - 0.480</p>\n<h3>Model 5 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['00f674', '8d43d9', '107712', '7b2e84', '575c60', '7f8949', '2874c5', '72e2c7']</p>\n<p>Train data: defog data, notype data<br>\nValidation data: defog data, notype data</p>\n<p>Train 45 minutes on TPU. Validation scores:<br>\nStartHesitation AP - [not used] Turn AP - 0.625 Walking AP - 0.238 mAP - 0.432<br>\nEvent AP - 0.800</p>\n<h3>Model 6 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n</code></pre>\n<p>Train data: defog data (about 85%)<br>\nValidation data: defog data (about 15%), notype data (100%)</p>\n<h3>Model 7 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Train data: defog data (100%)<br>\nValidation data: notype data (100%)</p>\n<p>Train 18 minutes on TPU. Validation scores:<br>\nStartHesitation AP - [not used] Turn AP - [not used] Walking AP - [not used] mAP - [not used]<br>\nEvent AP - 0.764</p>\n<h3>Model 8 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects<br>\n['12f8d1', '8c1f5e', '387ea0', 'c56629', '7da72f', '413532', 'd89567', 'ab3b2e', 'c83ff6', '056372']</p>\n<p>Train data: defog data, notype data<br>\nValidation data: defog data, notype data</p>\n<p>Train 28 minutes on TPU. Validation scores:<br>\nStartHesitation AP - [not used] Turn AP - 0.758 Walking AP - 0.221 mAP - 0.489<br>\nEvent AP - 0.744</p>\n<h1>Final models</h1>\n<p>Tdcsfog:  0.25 * Model 1 + 0.25 * Model 2 + 0.25 * Model 3 + 0.25 * Model 4</p>\n<p>Defog: 0.25 * Model 5 + 0.25 * Model 6 + 0.25 * Model 7 + 0.25 * Model 8</p>",
  "messages": [
    {
      "id": 2293457,
      "postDate": "2023-06-09T07:48:46.850Z",
      "content": "<p>Greetings to the Kaggle Community. In this message I want to tell you about my solution.</p>\n<p>Thanks to Kaggle for providing free GPU and TPU resources to everyone. On my graphics card (1050 Ti) I would not have achieved those results.<br>\nThanks to Google for the excellent tensorflow library.<br>\nAll of my work was done in Kaggle Notebooks and relies on TensorFlow capabilities.</p>\n<p>The key decisions that, in my opinion, led to a good result:</p>\n<ol>\n<li>Use a combination of transformer encoder and two BidirectionalLSTM layers.</li>\n<li>Use patches like VisualTransformer.</li>\n<li>Reduce the resolution of targets.</li>\n</ol>\n<p><em>How does it work?</em></p>\n<p>Suppose we have a tdcsfog sensor data series with AccV, AccML, AccAP columns and len of 5000.</p>\n<p>First, apply mean-std normalization to AccV, AccML, AccAP columns.</p>\n<pre><code> ():\n    mean = tf.math.reduce_mean(sample)\n    std = tf.math.reduce_std(sample)\n    sample = tf.math.divide_no_nan(sample-mean, std)\n\n     sample.numpy()\n</code></pre>\n<p>Then the series is zero-padded so that the final length is divisible by block_size = 15552  (or 12096 for defog). Now the series shape is (15552,  3). </p>\n<p>And create patches with the patch_size = 18 (or 14 for defog):</p>\n<pre><code>series \nseries = tf.reshape(series, shape=(CFG[] // CFG[], CFG[], )) \nseries = tf.reshape(series, shape=(CFG[] // CFG[], CFG[]*))  \n</code></pre>\n<p>Now the series shape is (864,  54). It's a model input.</p>\n<p>What to do with the StartHesitation, Turn, Walking data? Same, but apply tf.reduce_max at the end.</p>\n<pre><code>series_targets \nseries_targets = tf.reshape(series_targets, shape=(CFG[] // CFG[], CFG[], )) \nseries_targets = tf.transpose(series_targets, perm=[, , ]) \nseries_targets = tf.reduce_max(series_targets, axis=-) \n</code></pre>\n<p>Now the series shape is (864, 3). It's a model output.</p>\n<p>At the end, simply return the true resolution with tf.tile</p>\n<pre><code>predictions = model.predict(...) \npredictions = tf.expand_dims(predictions, axis=-) \npredictions = tf.transpose(predictions, perm=[, , , ]) \npredictions = tf.tile(predictions, multiples=[, , CFG[], ]) \npredictions = tf.reshape(predictions, shape=(predictions.shape[], predictions.shape[]*predictions.shape[], )) \n</code></pre>\n<h1>Details</h1>\n<p>Daily data, events.csv, subjects.csv, tasks.csv have never been used.</p>\n<p>Tdcsfog data is not used to train defog models. </p>\n<p>Defog data is not used to train tdcsfog models.</p>\n<p><em>Optimizer</em> </p>\n<pre><code>tf.keras.optimizers.Adam(learning_rate=Schedule(LEARNING_RATE, WARMUP_STEPS), beta_1=, beta_2=, epsilon=)\n</code></pre>\n<p><em>Loss function</em></p>\n<pre><code>\n\nce = tf.keras.losses.BinaryCrossentropy(reduction=)\n\n ():\n    loss = ce(tf.expand_dims(real[:, :, :], axis=-), tf.expand_dims(output, axis=-)) \n\n    mask = tf.math.multiply(real[:, :, ], real[:, :, ]) \n    mask = tf.cast(mask, dtype=loss.dtype)\n    mask = tf.expand_dims(mask, axis=-) \n    mask = tf.tile(mask, multiples=[, , ]) \n    loss *= mask \n\n     tf.reduce_sum(loss) / tf.reduce_sum(mask)\n</code></pre>\n<p><em>Model</em> </p>\n<pre><code>CFG = {: ,\n       : ,\n       : //,\n       : ,\n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\n\n\n (tf.keras.layers.Layer):\n     ():\n        ().__init__()\n\n        self.mha = tf.keras.layers.MultiHeadAttention(num_heads=CFG[], key_dim=CFG[], dropout=CFG[])\n\n        self.add = tf.keras.layers.Add()\n\n        self.layernorm = tf.keras.layers.LayerNormalization()\n\n        self.seq = tf.keras.Sequential([tf.keras.layers.Dense(CFG[], activation=),\n                                        tf.keras.layers.Dropout(CFG[]),\n                                        tf.keras.layers.Dense(CFG[]),\n                                        tf.keras.layers.Dropout(CFG[]),\n                                       ])\n\n     ():\n        attn_output = self.mha(query=x, key=x, value=x)\n        x = self.add([x, attn_output])\n        x = self.layernorm(x)\n        x = self.add([x, self.seq(x)])\n        x = self.layernorm(x)\n\n         x\n\n\n\n (tf.keras.Model):\n     ():\n        ().__init__()\n\n        self.first_linear = tf.keras.layers.Dense(CFG[])\n\n        self.add = tf.keras.layers.Add()\n\n        self.first_dropout = tf.keras.layers.Dropout(CFG[])\n\n        self.enc_layers = [EncoderLayer()  _  (CFG[])]\n\n        self.lstm_layers = [tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(CFG[], return_sequences=))  _  (CFG[])]\n\n        self.sequence_len = CFG[] // CFG[]\n        self.pos_encoding = tf.Variable(initial_value=tf.random.normal(shape=(, self.sequence_len, CFG[]), stddev=), trainable=)\n\n     (): \n        x = x /  \n        x = self.first_linear(x) \n\n         training: \n            random_pos_encoding = tf.roll(tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, , ]),\n                                          shift=tf.random.uniform(shape=(GPU_BATCH_SIZE,), minval=-self.sequence_len, maxval=, dtype=tf.int32),\n                                          axis=GPU_BATCH_SIZE * [],\n                                          )\n            x = self.add([x, random_pos_encoding])\n\n        : \n            x = self.add([x, tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, , ])])\n\n        x = self.first_dropout(x)\n\n         i  (CFG[]): x = self.enc_layers[i](x) \n         i  (CFG[]): x = self.lstm_layers[i](x) \n\n         x\n\n (tf.keras.Model):\n     ():\n        ().__init__()\n\n        self.encoder = FOGEncoder()\n        self.last_linear = tf.keras.layers.Dense()\n\n     (): \n        x = self.encoder(x) \n        x = self.last_linear(x) \n        x = tf.nn.sigmoid(x) \n\n         x\n</code></pre>\n<h1>Submission (Private Score 0.514, Public Score 0.527) consists of 8 models:</h1>\n<h3>Model 1 (tdcsfog model)</h3>\n<pre><code>CFG = {: , \n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE=\n</code></pre>\n<p>Validation subjects <br>\n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']</p>\n<p>Train 15 minutes on TPU. Validation scores:<br>\nStartHesitation AP - 0.462 Turn AP - 0.896 Walking AP - 0.470 mAP - 0.609</p>\n<h3>Model 2 (tdcsfog model)</h3>\n<pre><code>CFG = {: , \n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']</p>\n<p>Train 40 minutes on GPU. Validation scores:<br>\nStartHesitation AP - 0.481 Turn AP - 0.886 Walking AP - 0.437 mAP - 0.601</p>\n<h3>Model 3 (tdcsfog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['e39bc5', '516a67', 'af82b2', '4dc2f8', '743f4e', 'fa8764', 'a03db7', '51574c', '2d57c2']</p>\n<p>Train 11 minutes on TPU. Validation scores:<br>\nStartHesitation AP - 0.601 Turn AP - 0.857 Walking AP - 0.289 mAP - 0.582</p>\n<h3>Model 4 (tdcsfog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['5c0b8a', 'a03db7', '7fcee9', '2c98f7', '2a39f8', '4f13b4', 'af82b2', 'f686f0', '93f49f', '194d1d', '02bc69', '082f01']</p>\n<p>Train 13 minutes on TPU. Validation scores:<br>\nStartHesitation AP - 0.367 Turn AP - 0.879 Walking AP - 0.194 mAP - 0.480</p>\n<h3>Model 5 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects <br>\n['00f674', '8d43d9', '107712', '7b2e84', '575c60', '7f8949', '2874c5', '72e2c7']</p>\n<p>Train data: defog data, notype data<br>\nValidation data: defog data, notype data</p>\n<p>Train 45 minutes on TPU. Validation scores:<br>\nStartHesitation AP - [not used] Turn AP - 0.625 Walking AP - 0.238 mAP - 0.432<br>\nEvent AP - 0.800</p>\n<h3>Model 6 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n</code></pre>\n<p>Train data: defog data (about 85%)<br>\nValidation data: defog data (about 15%), notype data (100%)</p>\n<h3>Model 7 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Train data: defog data (100%)<br>\nValidation data: notype data (100%)</p>\n<p>Train 18 minutes on TPU. Validation scores:<br>\nStartHesitation AP - [not used] Turn AP - [not used] Walking AP - [not used] mAP - [not used]<br>\nEvent AP - 0.764</p>\n<h3>Model 8 (defog model)</h3>\n<pre><code>CFG = {: ,\n       : , \n       : //,\n       : , \n\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n       : ,\n      }\n\nLEARNING_RATE = /\nSTEPS_PER_EPOCH = \nWARMUP_STEPS = \nBATCH_SIZE = \n</code></pre>\n<p>Validation subjects<br>\n['12f8d1', '8c1f5e', '387ea0', 'c56629', '7da72f', '413532', 'd89567', 'ab3b2e', 'c83ff6', '056372']</p>\n<p>Train data: defog data, notype data<br>\nValidation data: defog data, notype data</p>\n<p>Train 28 minutes on TPU. Validation scores:<br>\nStartHesitation AP - [not used] Turn AP - 0.758 Walking AP - 0.221 mAP - 0.489<br>\nEvent AP - 0.744</p>\n<h1>Final models</h1>\n<p>Tdcsfog:  0.25 * Model 1 + 0.25 * Model 2 + 0.25 * Model 3 + 0.25 * Model 4</p>\n<p>Defog: 0.25 * Model 5 + 0.25 * Model 6 + 0.25 * Model 7 + 0.25 * Model 8</p>",
      "rawMarkdown": "Greetings to the Kaggle Community. In this message I want to tell you about my solution.\n\nThanks to Kaggle for providing free GPU and TPU resources to everyone. On my graphics card (1050 Ti) I would not have achieved those results.\nThanks to Google for the excellent tensorflow library.\nAll of my work was done in Kaggle Notebooks and relies on TensorFlow capabilities.\n\nThe key decisions that, in my opinion, led to a good result:\n1. Use a combination of transformer encoder and two BidirectionalLSTM layers.\n2. Use patches like VisualTransformer.\n3. Reduce the resolution of targets.\n\n*How does it work?*\n\nSuppose we have a tdcsfog sensor data series with AccV, AccML, AccAP columns and len of 5000.\n\nFirst, apply mean-std normalization to AccV, AccML, AccAP columns.\n\n```python\ndef sample_normalize(sample):\n\tmean = tf.math.reduce_mean(sample)\n\tstd = tf.math.reduce_std(sample)\n\tsample = tf.math.divide_no_nan(sample-mean, std)\n    \n\treturn sample.numpy()\n```\nThen the series is zero-padded so that the final length is divisible by block_size = 15552  (or 12096 for defog). Now the series shape is (15552,  3). \n\nAnd create patches with the patch_size = 18 (or 14 for defog):\n\n```python\nseries # Example shape (15552, 3)\nseries = tf.reshape(series, shape=(CFG['block_size'] // CFG['patch_size'], CFG['patch_size'], 3)) # Example shape (864, 18, 3)\nseries = tf.reshape(series, shape=(CFG['block_size'] // CFG['patch_size'], CFG['patch_size']*3))  # Example shape (864, 54)\n```\n\nNow the series shape is (864,  54). It's a model input.\n\nWhat to do with the StartHesitation, Turn, Walking data? Same, but apply tf.reduce_max at the end.\n\n```python\nseries_targets # Example shape (15552,  3)\nseries_targets = tf.reshape(series_targets, shape=(CFG['block_size'] // CFG['patch_size'], CFG['patch_size'], 3)) # Example shape (864, 18, 3)\nseries_targets = tf.transpose(series_targets, perm=[0, 2, 1]) # Example shape (864, 3, 18)\nseries_targets = tf.reduce_max(series_targets, axis=-1) # Example shape (864, 3)\n```\n\nNow the series shape is (864, 3). It's a model output.\n\nAt the end, simply return the true resolution with tf.tile\n\n```python\npredictions = model.predict(...) # Example shape (1, 864, 3)\npredictions = tf.expand_dims(predictions, axis=-1) # Example shape (1, 864, 3, 1)\npredictions = tf.transpose(predictions, perm=[0, 1, 3, 2]) # Example shape (1, 864, 1, 3)\npredictions = tf.tile(predictions, multiples=[1, 1, CFG['patch_size'], 1]) # Example shape (1, 864, 18, 3)\npredictions = tf.reshape(predictions, shape=(predictions.shape[0], predictions.shape[1]*predictions.shape[2], 3)) # Example shape (1, 15552, 3)\n```\n# Details\n\nDaily data, events.csv, subjects.csv, tasks.csv have never been used.\n\nTdcsfog data is not used to train defog models. \n\nDefog data is not used to train tdcsfog models.\n\n*Optimizer* \n\n```python\ntf.keras.optimizers.Adam(learning_rate=Schedule(LEARNING_RATE, WARMUP_STEPS), beta_1=0.9, beta_2=0.98, epsilon=1e-9)\n```\n\n*Loss function*\n\n```python\n'''\nloss_function args exp\n\nreal is a tensor with the shape (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], 5) where the last axis means:\n0 - StartHesitation\n1 - Turn\n2 - Walking\n3 - Valid\n4 - Mask\n\noutput is a tensor with the shape (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], 3) where the last axis means:\n0 - StartHesitation predicted\n1 - Turn predicted\n2 - Walking predicted\n\n'''\n\nce = tf.keras.losses.BinaryCrossentropy(reduction='none')\n\ndef loss_function(real, output, name='loss_function'):\n\tloss = ce(tf.expand_dims(real[:, :, 0:3], axis=-1), tf.expand_dims(output, axis=-1)) # Example shape (32, 864, 3)\n    \n\tmask = tf.math.multiply(real[:, :, 3], real[:, :, 4]) # Example shape (32, 864)\n\tmask = tf.cast(mask, dtype=loss.dtype)\n\tmask = tf.expand_dims(mask, axis=-1) # Example shape (32, 864, 1)\n\tmask = tf.tile(mask, multiples=[1, 1, 3]) # Example shape (32, 864, 3)\n\tloss *= mask # Example shape (32, 864, 3)\n\n\treturn tf.reduce_sum(loss) / tf.reduce_sum(mask)\n```\n*Model* \n\n```python\nCFG = {'TPU': 0,\n   \t'block_size': 15552,\n   \t'block_stride': 15552//16,\n   \t'patch_size': 18,\n  \t \n   \t'fog_model_dim': 320,\n   \t'fog_model_num_heads': 6,\n   \t'fog_model_num_encoder_layers': 5,\n   \t'fog_model_num_lstm_layers': 2,\n   \t'fog_model_first_dropout': 0.1,\n   \t'fog_model_encoder_dropout': 0.1,\n   \t'fog_model_mha_dropout': 0.0,\n  \t}\n\n'''\nThe transformer encoder layer\nFor more details, see https://arxiv.org/pdf/1706.03762.pdf [Attention Is All You Need]\n\n'''\n\nclass EncoderLayer(tf.keras.layers.Layer):\n\tdef __init__(self):\n    \tsuper().__init__()\n   \t \n    \tself.mha = tf.keras.layers.MultiHeadAttention(num_heads=CFG['fog_model_num_heads'], key_dim=CFG['fog_model_dim'], dropout=CFG['fog_model_mha_dropout'])\n   \t \n    \tself.add = tf.keras.layers.Add()\n   \t \n    \tself.layernorm = tf.keras.layers.LayerNormalization()\n   \t \n    \tself.seq = tf.keras.Sequential([tf.keras.layers.Dense(CFG['fog_model_dim'], activation='relu'),\n                                    \ttf.keras.layers.Dropout(CFG['fog_model_encoder_dropout']),\n                                    \ttf.keras.layers.Dense(CFG['fog_model_dim']),\n                                    \ttf.keras.layers.Dropout(CFG['fog_model_encoder_dropout']),\n                                   \t])\n   \t \n\tdef call(self, x):\n    \tattn_output = self.mha(query=x, key=x, value=x)\n    \tx = self.add([x, attn_output])\n    \tx = self.layernorm(x)\n    \tx = self.add([x, self.seq(x)])\n    \tx = self.layernorm(x)\n   \t \n    \treturn x\n    \n'''\nFOGEncoder is a combination of transformer encoder (D=320, H=6, L=5) and two BidirectionalLSTM layers\n\n'''\n\nclass FOGEncoder(tf.keras.Model):\n\tdef __init__(self):\n    \tsuper().__init__()\n   \t \n    \tself.first_linear = tf.keras.layers.Dense(CFG['fog_model_dim'])\n   \t \n    \tself.add = tf.keras.layers.Add()\n   \t \n    \tself.first_dropout = tf.keras.layers.Dropout(CFG['fog_model_first_dropout'])\n   \t \n    \tself.enc_layers = [EncoderLayer() for _ in range(CFG['fog_model_num_encoder_layers'])]\n   \t \n    \tself.lstm_layers = [tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(CFG['fog_model_dim'], return_sequences=True)) for _ in range(CFG['fog_model_num_lstm_layers'])]\n   \t \n    \tself.sequence_len = CFG['block_size'] // CFG['patch_size']\n    \tself.pos_encoding = tf.Variable(initial_value=tf.random.normal(shape=(1, self.sequence_len, CFG['fog_model_dim']), stddev=0.02), trainable=True)\n   \t \n\tdef call(self, x, training=None): # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['patch_size']*3), Example shape (4, 864, 54)\n    \tx = x / 25.0 # Normalization attempt in the segment [-1, 1]\n    \tx = self.first_linear(x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']), Example shape (4, 864, 320)\n     \t \n    \tif training: # augmentation by randomly roll of the position encoding tensor\n        \trandom_pos_encoding = tf.roll(tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, 1, 1]),\n                                      \tshift=tf.random.uniform(shape=(GPU_BATCH_SIZE,), minval=-self.sequence_len, maxval=0, dtype=tf.int32),\n                                      \taxis=GPU_BATCH_SIZE * [1],\n                                      \t)\n        \tx = self.add([x, random_pos_encoding])\n   \t \n    \telse: # without augmentation\n        \tx = self.add([x, tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, 1, 1])])\n       \t \n    \tx = self.first_dropout(x)\n   \t \n    \tfor i in range(CFG['fog_model_num_encoder_layers']): x = self.enc_layers[i](x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']), Example shape (4, 864, 320)\n    \tfor i in range(CFG['fog_model_num_lstm_layers']): x = self.lstm_layers[i](x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']*2), Example shape (4, 864, 640)\n       \t \n    \treturn x\n    \nclass FOGModel(tf.keras.Model):\n\tdef __init__(self):\n    \tsuper().__init__()\n   \t \n    \tself.encoder = FOGEncoder()\n    \tself.last_linear = tf.keras.layers.Dense(3)\n   \t \n\tdef call(self, x): # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['patch_size']*3), Example shape (4, 864, 54)\n    \tx = self.encoder(x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']*2), Example shape (4, 864, 640)\n    \tx = self.last_linear(x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], 3), Example shape (4, 864, 3)\n    \tx = tf.nn.sigmoid(x) # Sigmoid activation\n   \t \n    \treturn x\n\n```\n\n\n\n# Submission (Private Score 0.514, Public Score 0.527) consists of 8 models:\n\n### Model 1 (tdcsfog model)\n\n```python\nCFG = {'TPU': 1, \n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/38\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE=32\n```\n\nValidation subjects \n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']\n\nTrain 15 minutes on TPU. Validation scores:\nStartHesitation AP - 0.462 Turn AP - 0.896 Walking AP - 0.470 mAP - 0.609\n\n### Model 2 (tdcsfog model)\n\n```python\nCFG = {'TPU': 0, \n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 256,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 3,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/24\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE = 16\n```\n\nValidation subjects \n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']\n\nTrain 40 minutes on GPU. Validation scores:\nStartHesitation AP - 0.481 Turn AP - 0.886 Walking AP - 0.437 mAP - 0.601\n\n### Model 3 (tdcsfog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/48\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE = 32\n```\n\nValidation subjects \n['e39bc5', '516a67', 'af82b2', '4dc2f8', '743f4e', 'fa8764', 'a03db7', '51574c', '2d57c2']\n\nTrain 11 minutes on TPU. Validation scores:\nStartHesitation AP - 0.601 Turn AP - 0.857 Walking AP - 0.289 mAP - 0.582\n\n### Model 4 (tdcsfog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/38\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE = 32\n```\n\nValidation subjects \n['5c0b8a', 'a03db7', '7fcee9', '2c98f7', '2a39f8', '4f13b4', 'af82b2', 'f686f0', '93f49f', '194d1d', '02bc69', '082f01']\n\nTrain 13 minutes on TPU. Validation scores:\nStartHesitation AP - 0.367 Turn AP - 0.879 Walking AP - 0.194 mAP - 0.480\n\n### Model 5 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/62\nSTEPS_PER_EPOCH = 256\nWARMUP_STEPS = 256\nBATCH_SIZE = 32\n```\n\nValidation subjects \n['00f674', '8d43d9', '107712', '7b2e84', '575c60', '7f8949', '2874c5', '72e2c7']\n\nTrain data: defog data, notype data\nValidation data: defog data, notype data\n\nTrain 45 minutes on TPU. Validation scores:\nStartHesitation AP - [not used] Turn AP - 0.625 Walking AP - 0.238 mAP - 0.432\nEvent AP - 0.800\n\n### Model 6 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 5,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n```\n\nTrain data: defog data (about 85%)\nValidation data: defog data (about 15%), notype data (100%)\n\n### Model 7 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 4,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/24\nSTEPS_PER_EPOCH = 32\nWARMUP_STEPS = 64\nBATCH_SIZE = 128\n```\n\nTrain data: defog data (100%)\nValidation data: notype data (100%)\n\nTrain 18 minutes on TPU. Validation scores:\nStartHesitation AP - [not used] Turn AP - [not used] Walking AP - [not used] mAP - [not used]\nEvent AP - 0.764\n\n### Model 8 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/46\nSTEPS_PER_EPOCH = 256\nWARMUP_STEPS = 256\nBATCH_SIZE = 32\n```\n\nValidation subjects\n['12f8d1', '8c1f5e', '387ea0', 'c56629', '7da72f', '413532', 'd89567', 'ab3b2e', 'c83ff6', '056372']\n\nTrain data: defog data, notype data\nValidation data: defog data, notype data\n\nTrain 28 minutes on TPU. Validation scores:\nStartHesitation AP - [not used] Turn AP - 0.758 Walking AP - 0.221 mAP - 0.489\nEvent AP - 0.744\n\n# Final models\n\nTdcsfog:  0.25 * Model 1 + 0.25 * Model 2 + 0.25 * Model 3 + 0.25 * Model 4\n\nDefog: 0.25 * Model 5 + 0.25 * Model 6 + 0.25 * Model 7 + 0.25 * Model 8\n\n\n\n\n\n\n",
      "votes": 117
    },
    {
      "id": 2294914,
      "postDate": "2023-06-10T12:55:25.253Z",
      "content": "<p>Congratulations 🎉🎉🎉🎉<br>\nThe fact that I don't understand a single comment here shows that I have a long way to go. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations 🎉🎉🎉🎉\nThe fact that I don't understand a single comment here shows that I have a long way to go. Thanks for sharing.",
      "votes": 10
    },
    {
      "id": 2294662,
      "postDate": "2023-06-10T07:59:17.710Z",
      "content": "<p>Thank you for sharing your solution, time-series dimensionality reduction is the thing I haven't thought of, I have a question for you, -</p>\n<p>How do you perform inference for your models? <br>\nSuppose you've got the predictions for N tokens in the token sequence, which ones do you take? (e.g. all of them and then average predictions for overlapping tokens from different windows, or do you take only central token from each window?)</p>",
      "rawMarkdown": "Thank you for sharing your solution, time-series dimensionality reduction is the thing I haven't thought of, I have a question for you, -\n\nHow do you perform inference for your models? \nSuppose you've got the predictions for N tokens in the token sequence, which ones do you take? (e.g. all of them and then average predictions for overlapping tokens from different windows, or do you take only central token from each window?)",
      "votes": 3,
      "replies": [
        {
          "id": 2294727,
          "postDate": "2023-06-10T09:15:50.890Z",
          "content": "<p>Thanks. Yes, simple window roll ([0:15552], [972:16524], [1944:17496], … ) and average predictions for overlapping tokens from different windows.</p>",
          "rawMarkdown": "Thanks. Yes, simple window roll ([0:15552], [972:16524], [1944:17496], ... ) and average predictions for overlapping tokens from different windows.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2296202,
      "postDate": "2023-06-11T16:03:27.690Z",
      "content": "<p>I hope you will receive the prize money for 1st place.<br>\nIf Kaggle cannot deliver it to you because of the sanctions, then I would write to the Michael J. Fox Foundation if I were you, so that it would be possible for you somehow…</p>",
      "rawMarkdown": "I hope you will receive the prize money for 1st place.\nIf Kaggle cannot deliver it to you because of the sanctions, then I would write to the Michael J. Fox Foundation if I were you, so that it would be possible for you somehow...",
      "votes": 4,
      "replies": [
        {
          "id": 2296204,
          "postDate": "2023-06-11T16:09:15.433Z",
          "content": "<p>He fought with a 1050 Ti and some Kaggle free resources and achieved 1st place as a 1 person team and this is his 1st place and gold medal.<br>\nIt should be rewarded, not withheld regardless of the global situation.</p>",
          "rawMarkdown": "He fought with a 1050 Ti and some Kaggle free resources and achieved 1st place as a 1 person team and this is his 1st place and gold medal.\nIt should be rewarded, not withheld regardless of the global situation.",
          "votes": 9
        }
      ]
    },
    {
      "id": 2304787,
      "postDate": "2023-06-16T08:33:02.163Z",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a> A question, why did you go with the resolution reduction approach, was it to save memory or is there some intrinsic benefit?</p>",
      "rawMarkdown": "Congratulations! @baurzhanurazalinov A question, why did you go with the resolution reduction approach, was it to save memory or is there some intrinsic benefit?",
      "votes": 1,
      "replies": [
        {
          "id": 2304840,
          "postDate": "2023-06-16T09:21:17.333Z",
          "content": "<p>Thanks!</p>\n<ol>\n<li>No, it's not saving memory. Model works better with reduced resolution than with true resolution. I think deep learning models do badly on targets with complex structure (for example, 128Hz like in this competition), so they should be reduced.<br>\nIn addition, it allows you to use raw input data instead of using feature engineering.</li>\n</ol>",
          "rawMarkdown": "Thanks!\n\n1. No, it's not saving memory. Model works better with reduced resolution than with true resolution. I think deep learning models do badly on targets with complex structure (for example, 128Hz like in this competition), so they should be reduced.\nIn addition, it allows you to use raw input data instead of using feature engineering.\n",
          "votes": 2,
          "replies": [
            {
              "id": 2305539,
              "postDate": "2023-06-16T18:52:32.047Z",
              "content": "<p>By your second point, do you mean that the high resolution is too noisy to render the data effective? Essentially, you are denoising the data by reducing resolution? Not sure how feature engineering relates to the resolution here</p>",
              "rawMarkdown": "By your second point, do you mean that the high resolution is too noisy to render the data effective? Essentially, you are denoising the data by reducing resolution? Not sure how feature engineering relates to the resolution here",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2302346,
      "postDate": "2023-06-14T13:42:31.537Z",
      "content": "<p><a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a> Great job! Congratulations! 👋</p>",
      "rawMarkdown": "@baurzhanurazalinov Great job! Congratulations! 👋",
      "votes": 1
    },
    {
      "id": 2301748,
      "postDate": "2023-06-14T06:29:53.153Z",
      "content": "<p>congratulations for work sharing👍</p>",
      "rawMarkdown": "congratulations for work sharing👍",
      "votes": 1
    },
    {
      "id": 2297770,
      "postDate": "2023-06-12T19:44:19.020Z",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a>! And thanks for sharing your knowledge with the community!</p>\n<p>Impressive to see that you achieved such results using only AccV, AccML and AccAP data series…</p>\n<p>Curious to understand why you used block_size = 15552 (or 12096 for defog), and if you tried other values?</p>\n<p>Keep up with the great work💪🔥</p>",
      "rawMarkdown": "Congratulations, @baurzhanurazalinov! And thanks for sharing your knowledge with the community!\n\nImpressive to see that you achieved such results using only AccV, AccML and AccAP data series...\n\nCurious to understand why you used block_size = 15552 (or 12096 for defog), and if you tried other values?\n\nKeep up with the great work💪🔥",
      "votes": 1,
      "replies": [
        {
          "id": 2300019,
          "postDate": "2023-06-13T01:53:48.877Z",
          "content": "<p>Thanks!</p>\n<ol>\n<li>The model has two important parameters: patch size and sequence length.<br>\noptimal tdscfog patch size = 18<br>\noptimal defog patch size = 14<br>\noptimal sequence length = 864<br>\n864 * 18 = 15552 and 864 * 14 = 12096<br>\nUsing other parameters values lowers the validation scores.</li>\n</ol>",
          "rawMarkdown": "Thanks!\n1. The model has two important parameters: patch size and sequence length.\noptimal tdscfog patch size = 18\noptimal defog patch size = 14\noptimal sequence length = 864\n864 * 18 = 15552 and 864 * 14 = 12096\nUsing other parameters values lowers the validation scores.",
          "votes": 2,
          "replies": [
            {
              "id": 2302381,
              "postDate": "2023-06-14T14:02:33.927Z",
              "content": "<p>Really interesting… </p>\n<p>It seem's that using patches like VisualTransformers and BidirectionalLSTM layers enabled the model to capture both global and local information efficiently.</p>\n<p>Congratulations and thanks once again🙏</p>",
              "rawMarkdown": "Really interesting... \n\nIt seem's that using patches like VisualTransformers and BidirectionalLSTM layers enabled the model to capture both global and local information efficiently.\n\nCongratulations and thanks once again🙏"
            }
          ]
        }
      ]
    },
    {
      "id": 2297403,
      "postDate": "2023-06-12T14:53:05.240Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 2297309,
      "postDate": "2023-06-12T13:53:20.093Z",
      "content": "<p>Congrats! It's inspiring to see people come up with such innovative solutions.</p>",
      "rawMarkdown": "Congrats! It's inspiring to see people come up with such innovative solutions.",
      "votes": 1
    },
    {
      "id": 2295919,
      "postDate": "2023-06-11T11:49:12.293Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 2295693,
      "postDate": "2023-06-11T07:40:32.163Z",
      "content": "<p>Great job! Thank you for sharing! May I ask where I can see your complete code? If I could further learn, I will appreciate a lot!</p>",
      "rawMarkdown": "Great job! Thank you for sharing! May I ask where I can see your complete code? If I could further learn, I will appreciate a lot!",
      "votes": 1
    },
    {
      "id": 2293572,
      "postDate": "2023-06-09T09:37:35.720Z",
      "content": "<p>Congratulations, I'm gonna implement your solution and I hope to understand your insights. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations, I'm gonna implement your solution and I hope to understand your insights. Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 2293543,
      "postDate": "2023-06-09T08:54:26.037Z",
      "content": "<p>Amazing job well done! </p>",
      "rawMarkdown": "Amazing job well done! ",
      "votes": 1
    },
    {
      "id": 2297024,
      "postDate": "2023-06-12T10:26:54.420Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a>, for winning this competition! And thank you for sharing your great solution.<br>\nThe use of patches and the consolidation of targets are great ideas.<br>\nI have a few questions:</p>\n<ul>\n<li>Have you tried a transformer-only solution (without the LSTM part)?</li>\n<li>I'd never seen a random rolling position encoding built this way. Is this your own idea? Have you compared it to a regular position encoding?</li>\n<li>You have a parameter block_stride that indicates you create multiple samples from each subject (with overlap). But the batch size and samples per epoch are low. How many epochs did you use?</li>\n</ul>",
      "rawMarkdown": "Congratulations @baurzhanurazalinov, for winning this competition! And thank you for sharing your great solution.\nThe use of patches and the consolidation of targets are great ideas.\nI have a few questions:\n- Have you tried a transformer-only solution (without the LSTM part)?\n- I'd never seen a random rolling position encoding built this way. Is this your own idea? Have you compared it to a regular position encoding?\n- You have a parameter block_stride that indicates you create multiple samples from each subject (with overlap). But the batch size and samples per epoch are low. How many epochs did you use?",
      "votes": 2,
      "replies": [
        {
          "id": 2297150,
          "postDate": "2023-06-12T11:36:40.827Z",
          "content": "<p>Thanks!</p>\n<ol>\n<li>Yes. Without LSTM layers, the model does not work well. I think the main contribution of the transformer encoder is to classify events (StartHesitation, Turn or Walking?), and LSTM part provides continuous communication between neighboring tokens.</li>\n<li>Yes, it was my idea. I am not sure that adding positional encoding improves validation scores, but I have noticed that positional encoding does not harm the model, so I decided not to remove it. I suppose that requires additional tests.</li>\n<li>Yes. I got about 2000 samples for tdscfog dataset and about 25000 samples for defog dataset. Model weights do not exceed 72MB and work quickly. Average number of epochs for a well-trained model is 35 epochs.</li>\n</ol>",
          "rawMarkdown": "Thanks!\n1. Yes. Without LSTM layers, the model does not work well. I think the main contribution of the transformer encoder is to classify events (StartHesitation, Turn or Walking?), and LSTM part provides continuous communication between neighboring tokens.\n2. Yes, it was my idea. I am not sure that adding positional encoding improves validation scores, but I have noticed that positional encoding does not harm the model, so I decided not to remove it. I suppose that requires additional tests.\n3. Yes. I got about 2000 samples for tdscfog dataset and about 25000 samples for defog dataset. Model weights do not exceed 72MB and work quickly. Average number of epochs for a well-trained model is 35 epochs.\n\n",
          "votes": 1,
          "replies": [
            {
              "id": 2297191,
              "postDate": "2023-06-12T12:11:25.133Z",
              "content": "<p>Thanks for your quick reply, <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a>!</p>\n<blockquote>\n  <p>LSTM layers, the model does not work well. I think the main contribution of the transformer encoder is to classify events (StartHesitation, Turn or Walking?), and LSTM part provides continuous communication between neighboring tokens.</p>\n</blockquote>\n<p>This is a great finding. I wonder if the LSTM part provides the inductive bias that is missing in the Transformer. I'll start using this combination.</p>\n<blockquote>\n  <p>Yes, it was my idea.</p>\n</blockquote>\n<p>Excellent! I'd like to also test this approach.</p>",
              "rawMarkdown": "Thanks for your quick reply, @baurzhanurazalinov!\n\n> LSTM layers, the model does not work well. I think the main contribution of the transformer encoder is to classify events (StartHesitation, Turn or Walking?), and LSTM part provides continuous communication between neighboring tokens.\n\nThis is a great finding. I wonder if the LSTM part provides the inductive bias that is missing in the Transformer. I'll start using this combination.\n\n> Yes, it was my idea.\n\nExcellent! I'd like to also test this approach.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2295732,
      "postDate": "2023-06-11T08:08:54.037Z",
      "content": "<p>Amazing work! Very elegant and interesting!</p>",
      "rawMarkdown": "Amazing work! Very elegant and interesting!",
      "votes": 2
    },
    {
      "id": 2293493,
      "postDate": "2023-06-09T08:17:35.973Z",
      "content": "<p>How do you pick validation subjects? And also how do you use event model for the test set?</p>",
      "rawMarkdown": "How do you pick validation subjects? And also how do you use event model for the test set?",
      "votes": 2,
      "replies": [
        {
          "id": 2293537,
          "postDate": "2023-06-09T08:49:32.293Z",
          "content": "<ol>\n<li>I count the number of StartHesitation, Turn, Walking events in each subject with scipy.ndimage.label function. Then randomly select subset of subjects (about 15% subjects). If subset contains 20 - 30% StartHesitation events and 20 -  30% Walking events, I choose it.</li>\n<li>Event model is only used for training on defog + notype data (4-target model, in case of notype examples the first three targets are masked). During submission, the fourth target is ignored. Also, event AP is a good metric. </li>\n</ol>\n<pre><code>series[] = series[[, , ]].aggregate(, axis=)\nseries[] = series[[, , ]].aggregate(, axis=)\n</code></pre>",
          "rawMarkdown": "1. I count the number of StartHesitation, Turn, Walking events in each subject with scipy.ndimage.label function. Then randomly select subset of subjects (about 15% subjects). If subset contains 20 - 30% StartHesitation events and 20 -  30% Walking events, I choose it.\n2. Event model is only used for training on defog + notype data (4-target model, in case of notype examples the first three targets are masked). During submission, the fourth target is ignored. Also, event AP is a good metric. \n\n```python\nseries['Event'] = series[['StartHesitation', 'Turn', 'Walking']].aggregate('max', axis=1)\nseries['Event_prediction'] = series[['StartHesitation_prediction', 'Turn_prediction', 'Walking_prediction']].aggregate('max', axis=1)\n```",
          "votes": 5,
          "replies": [
            {
              "id": 2293552,
              "postDate": "2023-06-09T09:13:12.023Z",
              "content": "<p>Amazing validation trick👍</p>",
              "rawMarkdown": "Amazing validation trick👍",
              "votes": 1
            },
            {
              "id": 2295437,
              "postDate": "2023-06-11T00:54:07.063Z",
              "content": "<p><a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a> <br>\nCongratulations!! I have questions about your comment 1 and 2.</p>\n<ol>\n<li>Did you choose subset (contains 20-30% StartHesitation and Walking events) for both train and valid data?<br>\nAnd, didn't you cross validation?</li>\n<li>In your case, you use Event in addition to StartHesitation, Turn, and Walking. <br>\nAre you saying that the addition of the Event flag helped to improve the accuracy of the deep learning model?</li>\n</ol>",
              "rawMarkdown": "@baurzhanurazalinov \nCongratulations!! I have questions about your comment 1 and 2.\n1. Did you choose subset (contains 20-30% StartHesitation and Walking events) for both train and valid data?\n    And, didn't you cross validation?\n2. In your case, you use Event in addition to StartHesitation, Turn, and Walking. \n    Are you saying that the addition of the Event flag helped to improve the accuracy of the deep learning model?",
              "votes": 2
            },
            {
              "id": 2295508,
              "postDate": "2023-06-11T03:44:51.587Z",
              "content": "<p>Thanks.</p>\n<ol>\n<li>Valid data only (valid data = selected subset; train data = all data - selected subset)<br>\nCross validation is not used because of the low correlation between validation score and LB. I used three rules to select the trained model:<br>\nModel was trained for a long time.<br>\nModel at the end of long training gives good validation scores<br>\nModel gives good public LB score</li>\n<li>If you do not use Event flag, it is not possible to include notype data in the training process.<br>\nI thought that pseudo-labeling notype data would lower my validation scores, so I didn't use pseudo-labeling.</li>\n</ol>",
              "rawMarkdown": "Thanks.\n1. Valid data only (valid data = selected subset; train data = all data - selected subset)\nCross validation is not used because of the low correlation between validation score and LB. I used three rules to select the trained model:\nModel was trained for a long time.\nModel at the end of long training gives good validation scores\nModel gives good public LB score\n2. If you do not use Event flag, it is not possible to include notype data in the training process.\nI thought that pseudo-labeling notype data would lower my validation scores, so I didn't use pseudo-labeling.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2293485,
      "postDate": "2023-06-09T08:12:36.120Z",
      "content": "<p>Great solution. <br>\nI also use LSTM + transformer, and my sequence length is 5120. But I didn't use patches like you which seems very helpful to avoid overfitting.👍👍👍 </p>",
      "rawMarkdown": "Great solution. \nI also use LSTM + transformer, and my sequence length is 5120. But I didn't use patches like you which seems very helpful to avoid overfitting.👍👍👍 ",
      "votes": 2,
      "replies": [
        {
          "id": 2293507,
          "postDate": "2023-06-09T08:28:02.493Z",
          "content": "<p>Thanks. Patch size is the most important parameter of this model. Increasing or decreasing significantly degrades mAP. <br>\nIt can also be noticed that 18 * (100Hz / 128Hz) = 14.0625, where 18 - optimal tdcsfog patch size, 14 - optimal defog patch size, but the tdcsfog and defog models are independent.</p>",
          "rawMarkdown": "Thanks. Patch size is the most important parameter of this model. Increasing or decreasing significantly degrades mAP. \nIt can also be noticed that 18 * (100Hz / 128Hz) = 14.0625, where 18 - optimal tdcsfog patch size, 14 - optimal defog patch size, but the tdcsfog and defog models are independent.",
          "votes": 4
        }
      ]
    },
    {
      "id": 3168194,
      "postDate": "2025-04-02T08:09:56.757Z",
      "content": "<p>🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺</p>",
      "rawMarkdown": "🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺"
    },
    {
      "id": 2321083,
      "postDate": "2023-06-28T09:15:09.957Z",
      "content": "<p>Thank you for sharing, and congrats!<br>\nPadding each time-series to 15552 samples is very memory consuming.<br>\nHave you processed and trained one time-series at a time (e.g. using <code>model.train_on_batch</code>) or other techinques or am I wrong in something? :-)</p>",
      "rawMarkdown": "Thank you for sharing, and congrats!\nPadding each time-series to 15552 samples is very memory consuming.\nHave you processed and trained one time-series at a time (e.g. using `model.train_on_batch`) or other techinques or am I wrong in something? :-)",
      "replies": [
        {
          "id": 2321102,
          "postDate": "2023-06-28T09:33:44.123Z",
          "content": "<p>Thanks!</p>\n<p>All train data were loaded into RAM. No problems with memory.<br>\nYou can run my notebook if you want to make sure:<br>\n<a href=\"https://www.kaggle.com/code/baurzhanurazalinov/parkinson-s-freezing-tdcsfog-training-code\" target=\"_blank\">https://www.kaggle.com/code/baurzhanurazalinov/parkinson-s-freezing-tdcsfog-training-code</a></p>",
          "rawMarkdown": "Thanks!\n\nAll train data were loaded into RAM. No problems with memory.\nYou can run my notebook if you want to make sure:\nhttps://www.kaggle.com/code/baurzhanurazalinov/parkinson-s-freezing-tdcsfog-training-code",
          "votes": 1,
          "replies": [
            {
              "id": 2321129,
              "postDate": "2023-06-28T09:44:45.120Z",
              "content": "<p>Thank you I just realized there were the complete notebooks available!</p>",
              "rawMarkdown": "Thank you I just realized there were the complete notebooks available!"
            }
          ]
        }
      ]
    },
    {
      "id": 2609687,
      "postDate": "2024-01-19T16:28:01.803Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2609685,
      "postDate": "2024-01-19T16:26:58.233Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2391081,
      "postDate": "2023-08-15T01:26:33.630Z",
      "content": "<p>Thank you for sharing the solution!</p>",
      "rawMarkdown": "Thank you for sharing the solution!"
    },
    {
      "id": 2301359,
      "postDate": "2023-06-13T20:26:03.027Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!"
    },
    {
      "id": 2297033,
      "postDate": "2023-06-12T10:31:42.947Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing"
    }
  ],
  "comments": [
    {
      "id": 2294914,
      "author_name": "Ifeanyichukwu Nwobodo",
      "author_url": "",
      "post_date": "2023-06-10T12:55:25.253000",
      "content": "<p>Congratulations 🎉🎉🎉🎉<br>\nThe fact that I don't understand a single comment here shows that I have a long way to go. Thanks for sharing.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 2294662,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-06-10T07:59:17.710000",
      "content": "<p>Thank you for sharing your solution, time-series dimensionality reduction is the thing I haven't thought of, I have a question for you, -</p>\n<p>How do you perform inference for your models? <br>\nSuppose you've got the predictions for N tokens in the token sequence, which ones do you take? (e.g. all of them and then average predictions for overlapping tokens from different windows, or do you take only central token from each window?)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2294727,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-10T09:15:50.890000",
          "content": "<p>Thanks. Yes, simple window roll ([0:15552], [972:16524], [1944:17496], … ) and average predictions for overlapping tokens from different windows.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2296202,
      "author_name": "AmorfEvo",
      "author_url": "",
      "post_date": "2023-06-11T16:03:27.690000",
      "content": "<p>I hope you will receive the prize money for 1st place.<br>\nIf Kaggle cannot deliver it to you because of the sanctions, then I would write to the Michael J. Fox Foundation if I were you, so that it would be possible for you somehow…</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2296204,
          "author_name": "AmorfEvo",
          "author_url": "",
          "post_date": "2023-06-11T16:09:15.433000",
          "content": "<p>He fought with a 1050 Ti and some Kaggle free resources and achieved 1st place as a 1 person team and this is his 1st place and gold medal.<br>\nIt should be rewarded, not withheld regardless of the global situation.</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 2304787,
      "author_name": "Yijie Xu",
      "author_url": "",
      "post_date": "2023-06-16T08:33:02.163000",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a> A question, why did you go with the resolution reduction approach, was it to save memory or is there some intrinsic benefit?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2304840,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-16T09:21:17.333000",
          "content": "<p>Thanks!</p>\n<ol>\n<li>No, it's not saving memory. Model works better with reduced resolution than with true resolution. I think deep learning models do badly on targets with complex structure (for example, 128Hz like in this competition), so they should be reduced.<br>\nIn addition, it allows you to use raw input data instead of using feature engineering.</li>\n</ol>",
          "votes": 2,
          "replies": [
            {
              "id": 2305539,
              "author_name": "Yijie Xu",
              "author_url": "",
              "post_date": "2023-06-16T18:52:32.047000",
              "content": "<p>By your second point, do you mean that the high resolution is too noisy to render the data effective? Essentially, you are denoising the data by reducing resolution? Not sure how feature engineering relates to the resolution here</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2302346,
      "author_name": "Eugeniy Osetrov",
      "author_url": "",
      "post_date": "2023-06-14T13:42:31.537000",
      "content": "<p><a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a> Great job! Congratulations! 👋</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2301748,
      "author_name": "Bachar Acherif",
      "author_url": "",
      "post_date": "2023-06-14T06:29:53.153000",
      "content": "<p>congratulations for work sharing👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2297770,
      "author_name": "Vladimir Simões da Luz Junior",
      "author_url": "",
      "post_date": "2023-06-12T19:44:19.020000",
      "content": "<p>Congratulations, <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a>! And thanks for sharing your knowledge with the community!</p>\n<p>Impressive to see that you achieved such results using only AccV, AccML and AccAP data series…</p>\n<p>Curious to understand why you used block_size = 15552 (or 12096 for defog), and if you tried other values?</p>\n<p>Keep up with the great work💪🔥</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2300019,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-13T01:53:48.877000",
          "content": "<p>Thanks!</p>\n<ol>\n<li>The model has two important parameters: patch size and sequence length.<br>\noptimal tdscfog patch size = 18<br>\noptimal defog patch size = 14<br>\noptimal sequence length = 864<br>\n864 * 18 = 15552 and 864 * 14 = 12096<br>\nUsing other parameters values lowers the validation scores.</li>\n</ol>",
          "votes": 2,
          "replies": [
            {
              "id": 2302381,
              "author_name": "Vladimir Simões da Luz Junior",
              "author_url": "",
              "post_date": "2023-06-14T14:02:33.927000",
              "content": "<p>Really interesting… </p>\n<p>It seem's that using patches like VisualTransformers and BidirectionalLSTM layers enabled the model to capture both global and local information efficiently.</p>\n<p>Congratulations and thanks once again🙏</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2297403,
      "author_name": "cbrt343",
      "author_url": "",
      "post_date": "2023-06-12T14:53:05.240000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2297309,
      "author_name": "ARYAN PERSHAD",
      "author_url": "",
      "post_date": "2023-06-12T13:53:20.093000",
      "content": "<p>Congrats! It's inspiring to see people come up with such innovative solutions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2295919,
      "author_name": "Eishkaran Singh",
      "author_url": "",
      "post_date": "2023-06-11T11:49:12.293000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2295693,
      "author_name": "Lingduo Wang",
      "author_url": "",
      "post_date": "2023-06-11T07:40:32.163000",
      "content": "<p>Great job! Thank you for sharing! May I ask where I can see your complete code? If I could further learn, I will appreciate a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2293572,
      "author_name": "olivepicker",
      "author_url": "",
      "post_date": "2023-06-09T09:37:35.720000",
      "content": "<p>Congratulations, I'm gonna implement your solution and I hope to understand your insights. Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2293543,
      "author_name": "JM",
      "author_url": "",
      "post_date": "2023-06-09T08:54:26.037000",
      "content": "<p>Amazing job well done! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2297024,
      "author_name": "Ignacio Oguiza",
      "author_url": "",
      "post_date": "2023-06-12T10:26:54.420000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a>, for winning this competition! And thank you for sharing your great solution.<br>\nThe use of patches and the consolidation of targets are great ideas.<br>\nI have a few questions:</p>\n<ul>\n<li>Have you tried a transformer-only solution (without the LSTM part)?</li>\n<li>I'd never seen a random rolling position encoding built this way. Is this your own idea? Have you compared it to a regular position encoding?</li>\n<li>You have a parameter block_stride that indicates you create multiple samples from each subject (with overlap). But the batch size and samples per epoch are low. How many epochs did you use?</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2297150,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-12T11:36:40.827000",
          "content": "<p>Thanks!</p>\n<ol>\n<li>Yes. Without LSTM layers, the model does not work well. I think the main contribution of the transformer encoder is to classify events (StartHesitation, Turn or Walking?), and LSTM part provides continuous communication between neighboring tokens.</li>\n<li>Yes, it was my idea. I am not sure that adding positional encoding improves validation scores, but I have noticed that positional encoding does not harm the model, so I decided not to remove it. I suppose that requires additional tests.</li>\n<li>Yes. I got about 2000 samples for tdscfog dataset and about 25000 samples for defog dataset. Model weights do not exceed 72MB and work quickly. Average number of epochs for a well-trained model is 35 epochs.</li>\n</ol>",
          "votes": 1,
          "replies": [
            {
              "id": 2297191,
              "author_name": "Ignacio Oguiza",
              "author_url": "",
              "post_date": "2023-06-12T12:11:25.133000",
              "content": "<p>Thanks for your quick reply, <a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a>!</p>\n<blockquote>\n  <p>LSTM layers, the model does not work well. I think the main contribution of the transformer encoder is to classify events (StartHesitation, Turn or Walking?), and LSTM part provides continuous communication between neighboring tokens.</p>\n</blockquote>\n<p>This is a great finding. I wonder if the LSTM part provides the inductive bias that is missing in the Transformer. I'll start using this combination.</p>\n<blockquote>\n  <p>Yes, it was my idea.</p>\n</blockquote>\n<p>Excellent! I'd like to also test this approach.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2295732,
      "author_name": "Roy Segalz",
      "author_url": "",
      "post_date": "2023-06-11T08:08:54.037000",
      "content": "<p>Amazing work! Very elegant and interesting!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2293493,
      "author_name": "Shujun",
      "author_url": "",
      "post_date": "2023-06-09T08:17:35.973000",
      "content": "<p>How do you pick validation subjects? And also how do you use event model for the test set?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2293537,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-09T08:49:32.293000",
          "content": "<ol>\n<li>I count the number of StartHesitation, Turn, Walking events in each subject with scipy.ndimage.label function. Then randomly select subset of subjects (about 15% subjects). If subset contains 20 - 30% StartHesitation events and 20 -  30% Walking events, I choose it.</li>\n<li>Event model is only used for training on defog + notype data (4-target model, in case of notype examples the first three targets are masked). During submission, the fourth target is ignored. Also, event AP is a good metric. </li>\n</ol>\n<pre><code>series[] = series[[, , ]].aggregate(, axis=)\nseries[] = series[[, , ]].aggregate(, axis=)\n</code></pre>",
          "votes": 5,
          "replies": [
            {
              "id": 2293552,
              "author_name": "此般浅薄",
              "author_url": "",
              "post_date": "2023-06-09T09:13:12.023000",
              "content": "<p>Amazing validation trick👍</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2295437,
              "author_name": "HideBu",
              "author_url": "",
              "post_date": "2023-06-11T00:54:07.063000",
              "content": "<p><a href=\"https://www.kaggle.com/baurzhanurazalinov\" target=\"_blank\">@baurzhanurazalinov</a> <br>\nCongratulations!! I have questions about your comment 1 and 2.</p>\n<ol>\n<li>Did you choose subset (contains 20-30% StartHesitation and Walking events) for both train and valid data?<br>\nAnd, didn't you cross validation?</li>\n<li>In your case, you use Event in addition to StartHesitation, Turn, and Walking. <br>\nAre you saying that the addition of the Event flag helped to improve the accuracy of the deep learning model?</li>\n</ol>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2295508,
              "author_name": "Urazalinov Baurzhan",
              "author_url": "",
              "post_date": "2023-06-11T03:44:51.587000",
              "content": "<p>Thanks.</p>\n<ol>\n<li>Valid data only (valid data = selected subset; train data = all data - selected subset)<br>\nCross validation is not used because of the low correlation between validation score and LB. I used three rules to select the trained model:<br>\nModel was trained for a long time.<br>\nModel at the end of long training gives good validation scores<br>\nModel gives good public LB score</li>\n<li>If you do not use Event flag, it is not possible to include notype data in the training process.<br>\nI thought that pseudo-labeling notype data would lower my validation scores, so I didn't use pseudo-labeling.</li>\n</ol>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2293485,
      "author_name": "hyd",
      "author_url": "",
      "post_date": "2023-06-09T08:12:36.120000",
      "content": "<p>Great solution. <br>\nI also use LSTM + transformer, and my sequence length is 5120. But I didn't use patches like you which seems very helpful to avoid overfitting.👍👍👍 </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2293507,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-09T08:28:02.493000",
          "content": "<p>Thanks. Patch size is the most important parameter of this model. Increasing or decreasing significantly degrades mAP. <br>\nIt can also be noticed that 18 * (100Hz / 128Hz) = 14.0625, where 18 - optimal tdcsfog patch size, 14 - optimal defog patch size, but the tdcsfog and defog models are independent.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3168194,
      "author_name": "hip_174999",
      "author_url": "",
      "post_date": "2025-04-02T08:09:56.757000",
      "content": "<p>🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2321083,
      "author_name": "Alberto Annoni",
      "author_url": "",
      "post_date": "2023-06-28T09:15:09.957000",
      "content": "<p>Thank you for sharing, and congrats!<br>\nPadding each time-series to 15552 samples is very memory consuming.<br>\nHave you processed and trained one time-series at a time (e.g. using <code>model.train_on_batch</code>) or other techinques or am I wrong in something? :-)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2321102,
          "author_name": "Urazalinov Baurzhan",
          "author_url": "",
          "post_date": "2023-06-28T09:33:44.123000",
          "content": "<p>Thanks!</p>\n<p>All train data were loaded into RAM. No problems with memory.<br>\nYou can run my notebook if you want to make sure:<br>\n<a href=\"https://www.kaggle.com/code/baurzhanurazalinov/parkinson-s-freezing-tdcsfog-training-code\" target=\"_blank\">https://www.kaggle.com/code/baurzhanurazalinov/parkinson-s-freezing-tdcsfog-training-code</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2321129,
              "author_name": "Alberto Annoni",
              "author_url": "",
              "post_date": "2023-06-28T09:44:45.120000",
              "content": "<p>Thank you I just realized there were the complete notebooks available!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2609687,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-19T16:28:01.803000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2609685,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-19T16:26:58.233000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2391081,
      "author_name": "Ruiruiw",
      "author_url": "",
      "post_date": "2023-08-15T01:26:33.630000",
      "content": "<p>Thank you for sharing the solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2301359,
      "author_name": "Mark Gustetic",
      "author_url": "",
      "post_date": "2023-06-13T20:26:03.027000",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2297033,
      "author_name": "Pooja Chauhan",
      "author_url": "",
      "post_date": "2023-06-12T10:31:42.947000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2293457": "Greetings to the Kaggle Community. In this message I want to tell you about my solution.\n\nThanks to Kaggle for providing free GPU and TPU resources to everyone. On my graphics card (1050 Ti) I would not have achieved those results.\nThanks to Google for the excellent tensorflow library.\nAll of my work was done in Kaggle Notebooks and relies on TensorFlow capabilities.\n\nThe key decisions that, in my opinion, led to a good result:\n1. Use a combination of transformer encoder and two BidirectionalLSTM layers.\n2. Use patches like VisualTransformer.\n3. Reduce the resolution of targets.\n\n*How does it work?*\n\nSuppose we have a tdcsfog sensor data series with AccV, AccML, AccAP columns and len of 5000.\n\nFirst, apply mean-std normalization to AccV, AccML, AccAP columns.\n\n```python\ndef sample_normalize(sample):\n\tmean = tf.math.reduce_mean(sample)\n\tstd = tf.math.reduce_std(sample)\n\tsample = tf.math.divide_no_nan(sample-mean, std)\n    \n\treturn sample.numpy()\n```\nThen the series is zero-padded so that the final length is divisible by block_size = 15552  (or 12096 for defog). Now the series shape is (15552,  3). \n\nAnd create patches with the patch_size = 18 (or 14 for defog):\n\n```python\nseries # Example shape (15552, 3)\nseries = tf.reshape(series, shape=(CFG['block_size'] // CFG['patch_size'], CFG['patch_size'], 3)) # Example shape (864, 18, 3)\nseries = tf.reshape(series, shape=(CFG['block_size'] // CFG['patch_size'], CFG['patch_size']*3))  # Example shape (864, 54)\n```\n\nNow the series shape is (864,  54). It's a model input.\n\nWhat to do with the StartHesitation, Turn, Walking data? Same, but apply tf.reduce_max at the end.\n\n```python\nseries_targets # Example shape (15552,  3)\nseries_targets = tf.reshape(series_targets, shape=(CFG['block_size'] // CFG['patch_size'], CFG['patch_size'], 3)) # Example shape (864, 18, 3)\nseries_targets = tf.transpose(series_targets, perm=[0, 2, 1]) # Example shape (864, 3, 18)\nseries_targets = tf.reduce_max(series_targets, axis=-1) # Example shape (864, 3)\n```\n\nNow the series shape is (864, 3). It's a model output.\n\nAt the end, simply return the true resolution with tf.tile\n\n```python\npredictions = model.predict(...) # Example shape (1, 864, 3)\npredictions = tf.expand_dims(predictions, axis=-1) # Example shape (1, 864, 3, 1)\npredictions = tf.transpose(predictions, perm=[0, 1, 3, 2]) # Example shape (1, 864, 1, 3)\npredictions = tf.tile(predictions, multiples=[1, 1, CFG['patch_size'], 1]) # Example shape (1, 864, 18, 3)\npredictions = tf.reshape(predictions, shape=(predictions.shape[0], predictions.shape[1]*predictions.shape[2], 3)) # Example shape (1, 15552, 3)\n```\n# Details\n\nDaily data, events.csv, subjects.csv, tasks.csv have never been used.\n\nTdcsfog data is not used to train defog models. \n\nDefog data is not used to train tdcsfog models.\n\n*Optimizer* \n\n```python\ntf.keras.optimizers.Adam(learning_rate=Schedule(LEARNING_RATE, WARMUP_STEPS), beta_1=0.9, beta_2=0.98, epsilon=1e-9)\n```\n\n*Loss function*\n\n```python\n'''\nloss_function args exp\n\nreal is a tensor with the shape (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], 5) where the last axis means:\n0 - StartHesitation\n1 - Turn\n2 - Walking\n3 - Valid\n4 - Mask\n\noutput is a tensor with the shape (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], 3) where the last axis means:\n0 - StartHesitation predicted\n1 - Turn predicted\n2 - Walking predicted\n\n'''\n\nce = tf.keras.losses.BinaryCrossentropy(reduction='none')\n\ndef loss_function(real, output, name='loss_function'):\n\tloss = ce(tf.expand_dims(real[:, :, 0:3], axis=-1), tf.expand_dims(output, axis=-1)) # Example shape (32, 864, 3)\n    \n\tmask = tf.math.multiply(real[:, :, 3], real[:, :, 4]) # Example shape (32, 864)\n\tmask = tf.cast(mask, dtype=loss.dtype)\n\tmask = tf.expand_dims(mask, axis=-1) # Example shape (32, 864, 1)\n\tmask = tf.tile(mask, multiples=[1, 1, 3]) # Example shape (32, 864, 3)\n\tloss *= mask # Example shape (32, 864, 3)\n\n\treturn tf.reduce_sum(loss) / tf.reduce_sum(mask)\n```\n*Model* \n\n```python\nCFG = {'TPU': 0,\n   \t'block_size': 15552,\n   \t'block_stride': 15552//16,\n   \t'patch_size': 18,\n  \t \n   \t'fog_model_dim': 320,\n   \t'fog_model_num_heads': 6,\n   \t'fog_model_num_encoder_layers': 5,\n   \t'fog_model_num_lstm_layers': 2,\n   \t'fog_model_first_dropout': 0.1,\n   \t'fog_model_encoder_dropout': 0.1,\n   \t'fog_model_mha_dropout': 0.0,\n  \t}\n\n'''\nThe transformer encoder layer\nFor more details, see https://arxiv.org/pdf/1706.03762.pdf [Attention Is All You Need]\n\n'''\n\nclass EncoderLayer(tf.keras.layers.Layer):\n\tdef __init__(self):\n    \tsuper().__init__()\n   \t \n    \tself.mha = tf.keras.layers.MultiHeadAttention(num_heads=CFG['fog_model_num_heads'], key_dim=CFG['fog_model_dim'], dropout=CFG['fog_model_mha_dropout'])\n   \t \n    \tself.add = tf.keras.layers.Add()\n   \t \n    \tself.layernorm = tf.keras.layers.LayerNormalization()\n   \t \n    \tself.seq = tf.keras.Sequential([tf.keras.layers.Dense(CFG['fog_model_dim'], activation='relu'),\n                                    \ttf.keras.layers.Dropout(CFG['fog_model_encoder_dropout']),\n                                    \ttf.keras.layers.Dense(CFG['fog_model_dim']),\n                                    \ttf.keras.layers.Dropout(CFG['fog_model_encoder_dropout']),\n                                   \t])\n   \t \n\tdef call(self, x):\n    \tattn_output = self.mha(query=x, key=x, value=x)\n    \tx = self.add([x, attn_output])\n    \tx = self.layernorm(x)\n    \tx = self.add([x, self.seq(x)])\n    \tx = self.layernorm(x)\n   \t \n    \treturn x\n    \n'''\nFOGEncoder is a combination of transformer encoder (D=320, H=6, L=5) and two BidirectionalLSTM layers\n\n'''\n\nclass FOGEncoder(tf.keras.Model):\n\tdef __init__(self):\n    \tsuper().__init__()\n   \t \n    \tself.first_linear = tf.keras.layers.Dense(CFG['fog_model_dim'])\n   \t \n    \tself.add = tf.keras.layers.Add()\n   \t \n    \tself.first_dropout = tf.keras.layers.Dropout(CFG['fog_model_first_dropout'])\n   \t \n    \tself.enc_layers = [EncoderLayer() for _ in range(CFG['fog_model_num_encoder_layers'])]\n   \t \n    \tself.lstm_layers = [tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(CFG['fog_model_dim'], return_sequences=True)) for _ in range(CFG['fog_model_num_lstm_layers'])]\n   \t \n    \tself.sequence_len = CFG['block_size'] // CFG['patch_size']\n    \tself.pos_encoding = tf.Variable(initial_value=tf.random.normal(shape=(1, self.sequence_len, CFG['fog_model_dim']), stddev=0.02), trainable=True)\n   \t \n\tdef call(self, x, training=None): # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['patch_size']*3), Example shape (4, 864, 54)\n    \tx = x / 25.0 # Normalization attempt in the segment [-1, 1]\n    \tx = self.first_linear(x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']), Example shape (4, 864, 320)\n     \t \n    \tif training: # augmentation by randomly roll of the position encoding tensor\n        \trandom_pos_encoding = tf.roll(tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, 1, 1]),\n                                      \tshift=tf.random.uniform(shape=(GPU_BATCH_SIZE,), minval=-self.sequence_len, maxval=0, dtype=tf.int32),\n                                      \taxis=GPU_BATCH_SIZE * [1],\n                                      \t)\n        \tx = self.add([x, random_pos_encoding])\n   \t \n    \telse: # without augmentation\n        \tx = self.add([x, tf.tile(self.pos_encoding, multiples=[GPU_BATCH_SIZE, 1, 1])])\n       \t \n    \tx = self.first_dropout(x)\n   \t \n    \tfor i in range(CFG['fog_model_num_encoder_layers']): x = self.enc_layers[i](x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']), Example shape (4, 864, 320)\n    \tfor i in range(CFG['fog_model_num_lstm_layers']): x = self.lstm_layers[i](x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']*2), Example shape (4, 864, 640)\n       \t \n    \treturn x\n    \nclass FOGModel(tf.keras.Model):\n\tdef __init__(self):\n    \tsuper().__init__()\n   \t \n    \tself.encoder = FOGEncoder()\n    \tself.last_linear = tf.keras.layers.Dense(3)\n   \t \n\tdef call(self, x): # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['patch_size']*3), Example shape (4, 864, 54)\n    \tx = self.encoder(x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], CFG['fog_model_dim']*2), Example shape (4, 864, 640)\n    \tx = self.last_linear(x) # (GPU_BATCH_SIZE, CFG['block_size'] // CFG['patch_size'], 3), Example shape (4, 864, 3)\n    \tx = tf.nn.sigmoid(x) # Sigmoid activation\n   \t \n    \treturn x\n\n```\n\n\n\n# Submission (Private Score 0.514, Public Score 0.527) consists of 8 models:\n\n### Model 1 (tdcsfog model)\n\n```python\nCFG = {'TPU': 1, \n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/38\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE=32\n```\n\nValidation subjects \n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']\n\nTrain 15 minutes on TPU. Validation scores:\nStartHesitation AP - 0.462 Turn AP - 0.896 Walking AP - 0.470 mAP - 0.609\n\n### Model 2 (tdcsfog model)\n\n```python\nCFG = {'TPU': 0, \n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 256,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 3,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/24\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE = 16\n```\n\nValidation subjects \n['07285e', '220a17', '54ee6e', '312788', '24a59d', '4bb5d0', '48fd62', '79011a', '7688c1']\n\nTrain 40 minutes on GPU. Validation scores:\nStartHesitation AP - 0.481 Turn AP - 0.886 Walking AP - 0.437 mAP - 0.601\n\n### Model 3 (tdcsfog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/48\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE = 32\n```\n\nValidation subjects \n['e39bc5', '516a67', 'af82b2', '4dc2f8', '743f4e', 'fa8764', 'a03db7', '51574c', '2d57c2']\n\nTrain 11 minutes on TPU. Validation scores:\nStartHesitation AP - 0.601 Turn AP - 0.857 Walking AP - 0.289 mAP - 0.582\n\n### Model 4 (tdcsfog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 15552, \n       'block_stride': 15552//16,\n       'patch_size': 18, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/38\nSTEPS_PER_EPOCH = 64\nWARMUP_STEPS = 64\nBATCH_SIZE = 32\n```\n\nValidation subjects \n['5c0b8a', 'a03db7', '7fcee9', '2c98f7', '2a39f8', '4f13b4', 'af82b2', 'f686f0', '93f49f', '194d1d', '02bc69', '082f01']\n\nTrain 13 minutes on TPU. Validation scores:\nStartHesitation AP - 0.367 Turn AP - 0.879 Walking AP - 0.194 mAP - 0.480\n\n### Model 5 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/62\nSTEPS_PER_EPOCH = 256\nWARMUP_STEPS = 256\nBATCH_SIZE = 32\n```\n\nValidation subjects \n['00f674', '8d43d9', '107712', '7b2e84', '575c60', '7f8949', '2874c5', '72e2c7']\n\nTrain data: defog data, notype data\nValidation data: defog data, notype data\n\nTrain 45 minutes on TPU. Validation scores:\nStartHesitation AP - [not used] Turn AP - 0.625 Walking AP - 0.238 mAP - 0.432\nEvent AP - 0.800\n\n### Model 6 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 5,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n```\n\nTrain data: defog data (about 85%)\nValidation data: defog data (about 15%), notype data (100%)\n\n### Model 7 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 4,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/24\nSTEPS_PER_EPOCH = 32\nWARMUP_STEPS = 64\nBATCH_SIZE = 128\n```\n\nTrain data: defog data (100%)\nValidation data: notype data (100%)\n\nTrain 18 minutes on TPU. Validation scores:\nStartHesitation AP - [not used] Turn AP - [not used] Walking AP - [not used] mAP - [not used]\nEvent AP - 0.764\n\n### Model 8 (defog model)\n\n```python\nCFG = {'TPU': 1,\n       'block_size': 12096, \n       'block_stride': 12096//16,\n       'patch_size': 14, \n       \n       'fog_model_dim': 320,\n       'fog_model_num_heads': 6,\n       'fog_model_num_encoder_layers': 5,\n       'fog_model_num_lstm_layers': 2,\n       'fog_model_first_dropout': 0.1,\n       'fog_model_encoder_dropout': 0.1,\n       'fog_model_mha_dropout': 0.0,\n      }\n\nLEARNING_RATE = 0.01/46\nSTEPS_PER_EPOCH = 256\nWARMUP_STEPS = 256\nBATCH_SIZE = 32\n```\n\nValidation subjects\n['12f8d1', '8c1f5e', '387ea0', 'c56629', '7da72f', '413532', 'd89567', 'ab3b2e', 'c83ff6', '056372']\n\nTrain data: defog data, notype data\nValidation data: defog data, notype data\n\nTrain 28 minutes on TPU. Validation scores:\nStartHesitation AP - [not used] Turn AP - 0.758 Walking AP - 0.221 mAP - 0.489\nEvent AP - 0.744\n\n# Final models\n\nTdcsfog:  0.25 * Model 1 + 0.25 * Model 2 + 0.25 * Model 3 + 0.25 * Model 4\n\nDefog: 0.25 * Model 5 + 0.25 * Model 6 + 0.25 * Model 7 + 0.25 * Model 8\n\n\n\n\n\n\n",
    "2294914": "Congratulations 🎉🎉🎉🎉\nThe fact that I don't understand a single comment here shows that I have a long way to go. Thanks for sharing.",
    "2294662": "Thank you for sharing your solution, time-series dimensionality reduction is the thing I haven't thought of, I have a question for you, -\n\nHow do you perform inference for your models? \nSuppose you've got the predictions for N tokens in the token sequence, which ones do you take? (e.g. all of them and then average predictions for overlapping tokens from different windows, or do you take only central token from each window?)",
    "2296202": "I hope you will receive the prize money for 1st place.\nIf Kaggle cannot deliver it to you because of the sanctions, then I would write to the Michael J. Fox Foundation if I were you, so that it would be possible for you somehow...",
    "2304787": "Congratulations! @baurzhanurazalinov A question, why did you go with the resolution reduction approach, was it to save memory or is there some intrinsic benefit?",
    "2302346": "@baurzhanurazalinov Great job! Congratulations! 👋",
    "2301748": "congratulations for work sharing👍",
    "2297770": "Congratulations, @baurzhanurazalinov! And thanks for sharing your knowledge with the community!\n\nImpressive to see that you achieved such results using only AccV, AccML and AccAP data series...\n\nCurious to understand why you used block_size = 15552 (or 12096 for defog), and if you tried other values?\n\nKeep up with the great work💪🔥",
    "2297403": "Congratulations!",
    "2297309": "Congrats! It's inspiring to see people come up with such innovative solutions.",
    "2295919": "Congratulations",
    "2295693": "Great job! Thank you for sharing! May I ask where I can see your complete code? If I could further learn, I will appreciate a lot!",
    "2293572": "Congratulations, I'm gonna implement your solution and I hope to understand your insights. Thanks for sharing.",
    "2293543": "Amazing job well done! ",
    "2297024": "Congratulations @baurzhanurazalinov, for winning this competition! And thank you for sharing your great solution.\nThe use of patches and the consolidation of targets are great ideas.\nI have a few questions:\n- Have you tried a transformer-only solution (without the LSTM part)?\n- I'd never seen a random rolling position encoding built this way. Is this your own idea? Have you compared it to a regular position encoding?\n- You have a parameter block_stride that indicates you create multiple samples from each subject (with overlap). But the batch size and samples per epoch are low. How many epochs did you use?",
    "2295732": "Amazing work! Very elegant and interesting!",
    "2293493": "How do you pick validation subjects? And also how do you use event model for the test set?",
    "2293485": "Great solution. \nI also use LSTM + transformer, and my sequence length is 5120. But I didn't use patches like you which seems very helpful to avoid overfitting.👍👍👍 ",
    "3168194": "🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺🐮🍺",
    "2321083": "Thank you for sharing, and congrats!\nPadding each time-series to 15552 samples is very memory consuming.\nHave you processed and trained one time-series at a time (e.g. using `model.train_on_batch`) or other techinques or am I wrong in something? :-)",
    "2609687": "",
    "2609685": "",
    "2391081": "Thank you for sharing the solution!",
    "2301359": "Congratulations and thanks for sharing!",
    "2297033": "Thank you for sharing"
  }
}