{
  "id": 543289,
  "title": "Did anyone find making lag features useful?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543289",
  "author_name": "",
  "post_date": "2024-10-29T17:09:47.328845100Z",
  "votes": 3,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I was wondering if anyone made lag features that helped his LB score improve. So far I tried do some feature engineering and create lag features but it always led to far worse score. I wonder if lags are useful or I should avoid them in this competition.</p>\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": "3031421",
      "postDate": "10/29/2024 17:09:47",
      "content": "<p>I was wondering if anyone made lag features that helped his LB score improve. So far I tried do some feature engineering and create lag features but it always led to far worse score. I wonder if lags are useful or I should avoid them in this competition.</p>\n<p>Thanks in advance!</p>",
      "rawMarkdown": "I was wondering if anyone made lag features that helped his LB score improve. So far I tried do some feature engineering and create lag features but it always led to far worse score. I wonder if lags are useful or I should avoid them in this competition.\n\nThanks in advance!",
      "votes": null
    },
    {
      "id": "3031437",
      "postDate": "10/29/2024 17:25:36",
      "content": "<p>Many features are lagged (or \"semi-lagged\") from each other, so lagging all features will not be helpful (only cause redundancy). Some features are not suitable for making it lagged. You need to properly choose the right feature to lag the right timesteps.</p>",
      "rawMarkdown": "Many features are lagged (or \"semi-lagged\") from each other, so lagging all features will not be helpful (only cause redundancy). Some features are not suitable for making it lagged. You need to properly choose the right feature to lag the right timesteps.",
      "votes": null
    },
    {
      "id": "3031453",
      "postDate": "10/29/2024 17:41:34",
      "content": "<p>Yes, so far I tried to lag only the target column (responder_6), but it didn't lead to good results (dropped my 0.43 LB score significantly to -0.024!), which was strange to me because to my logic and understanding the previous values of the variable we are trying to predict are usually very important in time-series problems.</p>",
      "rawMarkdown": "Yes, so far I tried to lag only the target column (responder_6), but it didn't lead to good results (dropped my 0.43 LB score significantly to -0.024!), which was strange to me because to my logic and understanding the previous values of the variable we are trying to predict are usually very important in time-series problems.",
      "votes": null
    },
    {
      "id": "3031468",
      "postDate": "10/29/2024 18:02:16",
      "content": "<p>this drop in LB score is insane!! you must have introduced a bug to your training  or predicting code while doing feature engineering  </p>",
      "rawMarkdown": "this drop in LB score is insane!! you must have introduced a bug to your training  or predicting code while doing feature engineering",
      "votes": null
    },
    {
      "id": "3031494",
      "postDate": "10/29/2024 18:32:31",
      "content": "<p>How did you implement the lag? The drop is quite crazy. I suspect some wrong doing here. </p>",
      "rawMarkdown": "How did you implement the lag? The drop is quite crazy. I suspect some wrong doing here.",
      "votes": null
    },
    {
      "id": "3032013",
      "postDate": "10/30/2024 13:00:59",
      "content": "<p>You can check out this notebook for a solid example on using <code>lag_1</code> on responders in inference. </p>\n<p><a href=\"https://www.kaggle.com/code/motono0223/js24-inference-gbdt-with-lags-singlemodel\" target=\"_blank\">https://www.kaggle.com/code/motono0223/js24-inference-gbdt-with-lags-singlemodel</a></p>",
      "rawMarkdown": "You can check out this notebook for a solid example on using `lag_1` on responders in inference. \n\nhttps://www.kaggle.com/code/motono0223/js24-inference-gbdt-with-lags-singlemodel",
      "votes": null
    },
    {
      "id": "3032264",
      "postDate": "10/30/2024 18:15:42",
      "content": "<p>i also face the same problem</p>",
      "rawMarkdown": "i also face the same problem",
      "votes": null
    },
    {
      "id": "3032424",
      "postDate": "10/30/2024 21:45:02",
      "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> I created the lags feature with a simple shift:<br>\npl.col('responder_6').shift(1).over(partition_by='symbol_id',order_by=['date_id', 'time_id']).alias('lag_1time responder_6')</p>\n<p>But I suspect you are right and I did something wrong in using the API. I got a feeling that I didn't understand the API completely. Only today I realized that the actual test data is served from the API in batches and not all in one go, so I think the loading part of the lag dataframe gone wrong.<br>\nDo we know if it is guaranteed that each batch from API has records from the same date_id and time_id? And it is guaranteed that the batches are served chronologically?</p>\n<p>I'll try to debug my code a bit more to understand it better.<br>\nAppreciate your help!</p>",
      "rawMarkdown": "shiyili I created the lags feature with a simple shift:\npl.col('responder_6').shift(1).over(partition_by='symbol_id',order_by=['date_id', 'time_id']).alias('lag_1time responder_6')\n\nBut I suspect you are right and I did something wrong in using the API. I got a feeling that I didn't understand the API completely. Only today I realized that the actual test data is served from the API in batches and not all in one go, so I think the loading part of the lag dataframe gone wrong.\nDo we know if it is guaranteed that each batch from API has records from the same date_id and time_id? And it is guaranteed that the batches are served chronologically?\n\nI'll try to debug my code a bit more to understand it better.\nAppreciate your help!",
      "votes": null
    },
    {
      "id": "3032428",
      "postDate": "10/30/2024 21:48:00",
      "content": "<p>Thank you very much! I'll check it out. I think I misunderstood the API part and I believe this notebook will clear up my confusion</p>",
      "rawMarkdown": "Thank you very much! I'll check it out. I think I misunderstood the API part and I believe this notebook will clear up my confusion",
      "votes": null
    },
    {
      "id": "3032468",
      "postDate": "10/30/2024 23:27:20",
      "content": "<p>Your shit(1).over() seems to be wrong. You don't lag the responders by 1 time-step, instead you should lag them by 1 day.</p>\n<p>If you do the following plot, you will find the raw responder_6 and the lagged are not matching.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F3dd6c0a2b7c65d7d8ca2ed78aa894621%2FScreenshot%202024-10-31%20at%2000.25.19.png?generation=1730330724207812&amp;alt=media\" alt=\"\"></p>\n<p>You should instead do:</p>\n<pre><code>df_all = pl.read_parquet(train_partitions, =[, , , ]).with_columns(\n        pl.col().shift().([, ]).()\n    )\n</code></pre>\n<p>which will generate matching series:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F54d4464f5be7cabaef5d0c6866d1a395%2FScreenshot%202024-10-31%20at%2000.26.39.png?generation=1730330808063141&amp;alt=media\" alt=\"\"></p>\n<p>Plotting code: </p>\n<pre><code>data_dir = Path(\"./data\")\ntrain_partitions = [data_dir / f\"train.parquet/partition_id={i}/part-0.parquet\"  i  range(, )]\n\ndf_all = pl.read_parquet(train_partitions, =[, , , ]).with_columns(\n        pl.col().shift().(partition_by=,order_by=[, ]).()\n    )\n\np1 = df_all.((pl.col()==)&amp;(pl.col()==)).([, ]).to_pandas()\np2 = df_all.((pl.col()==)&amp;(pl.col()==)).([, ]).to_pandas()\n\nplt.figure(figsize=(, ))\nplt.plot(p1[], p1[], label=)\nplt.plot(p2[], p2[], label=)\n\nplt.legend()\nplt.()\n</code></pre>",
      "rawMarkdown": "Your shit(1).over() seems to be wrong. You don't lag the responders by 1 time-step, instead you should lag them by 1 day.\n\nIf you do the following plot, you will find the raw responder_6 and the lagged are not matching.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F3dd6c0a2b7c65d7d8ca2ed78aa894621%2FScreenshot%202024-10-31%20at%2000.25.19.png?generation=1730330724207812&alt=media)\n\nYou should instead do:\n\n```\ndf_all = pl.read_parquet(train_partitions, columns=['date_id', 'time_id', 'symbol_id', 'responder_6']).with_columns(\n        pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('lag_1time responder_6')\n    )\n```\n\nwhich will generate matching series:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F54d4464f5be7cabaef5d0c6866d1a395%2FScreenshot%202024-10-31%20at%2000.26.39.png?generation=1730330808063141&alt=media)\n\nPlotting code: \n\n```\ndata_dir = Path(\"./data\")\ntrain_partitions = [data_dir / f\"train.parquet/partition_id={i}/part-0.parquet\" for i in range(9, 10)]\n\ndf_all = pl.read_parquet(train_partitions, columns=['date_id', 'time_id', 'symbol_id', 'responder_6']).with_columns(\n        pl.col('responder_6').shift(1).over(partition_by='symbol_id',order_by=['date_id', 'time_id']).alias('lag_1time responder_6')\n    )\n\np1 = df_all.filter((pl.col('date_id')==1530)&(pl.col('symbol_id')==0)).select(['time_id', 'responder_6']).to_pandas()\np2 = df_all.filter((pl.col('date_id')==1531)&(pl.col('symbol_id')==0)).select(['time_id', 'lag_1time responder_6']).to_pandas()\n\nplt.figure(figsize=(12, 3))\nplt.plot(p1['time_id'], p1['responder_6'], label='responder_6')\nplt.plot(p2['time_id'], p2['lag_1time responder_6'], label='lag_1time responder_6')\n\nplt.legend()\nplt.show()\n```",
      "votes": null
    },
    {
      "id": "3032495",
      "postDate": "10/31/2024 00:49:34",
      "content": "<p>Oh wow, seems like a severe mistake. Thank you for pointing it out!<br>\nCould you explain me the reason why we don't lag by time step, and why it's better to lag by date?<br>\nI want to have a deep understanding in time-series best practices.</p>",
      "rawMarkdown": "Oh wow, seems like a severe mistake. Thank you for pointing it out!\nCould you explain me the reason why we don't lag by time step, and why it's better to lag by date?\nI want to have a deep understanding in time-series best practices.",
      "votes": null
    },
    {
      "id": "3032504",
      "postDate": "10/31/2024 01:37:05",
      "content": "<p>It’s not better to lag by date, it’s just how lags is provided in the evaluation api. </p>",
      "rawMarkdown": "It’s not better to lag by date, it’s just how lags is provided in the evaluation api.",
      "votes": null
    },
    {
      "id": "3032703",
      "postDate": "10/31/2024 08:52:11",
      "content": "<p>It’s not a best practice. You do so to bc this is how the api serves data </p>",
      "rawMarkdown": "It’s not a best practice. You do so to bc this is how the api serves data",
      "votes": null
    },
    {
      "id": "3032944",
      "postDate": "10/31/2024 15:16:04",
      "content": "<p>I see. Thank you all very much for the help and the clarifications! I feel like I learnt a lot now.</p>",
      "rawMarkdown": "I see. Thank you all very much for the help and the clarifications! I feel like I learnt a lot now.",
      "votes": null
    },
    {
      "id": "3081039",
      "postDate": "12/26/2024 06:15:10",
      "content": "<p>I added extra feature_00_lags to features_78_lags besides the original features, but it seems it doesn work, have you tried it?</p>",
      "rawMarkdown": "I added extra feature_00_lags to features_78_lags besides the original features, but it seems it doesn work, have you tried it?",
      "votes": null
    },
    {
      "id": "3084180",
      "postDate": "12/30/2024 13:56:48",
      "content": "<p>If I made it correctly I tested with 4 feature lags using tree models, it not change much offline and got much worse LB.</p>",
      "rawMarkdown": "If I made it correctly I tested with 4 feature lags using tree models, it not change much offline and got much worse LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3031437,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/29/2024 17:25:36",
      "content": "<p>Many features are lagged (or \"semi-lagged\") from each other, so lagging all features will not be helpful (only cause redundancy). Some features are not suitable for making it lagged. You need to properly choose the right feature to lag the right timesteps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3031453,
          "author_name": "dinezra11",
          "author_url": "",
          "post_date": "10/29/2024 17:41:34",
          "content": "<p>Yes, so far I tried to lag only the target column (responder_6), but it didn't lead to good results (dropped my 0.43 LB score significantly to -0.024!), which was strange to me because to my logic and understanding the previous values of the variable we are trying to predict are usually very important in time-series problems.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3031468,
              "author_name": "aymanallawi",
              "author_url": "",
              "post_date": "10/29/2024 18:02:16",
              "content": "<p>this drop in LB score is insane!! you must have introduced a bug to your training  or predicting code while doing feature engineering  </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3031494,
              "author_name": "shiyili",
              "author_url": "",
              "post_date": "10/29/2024 18:32:31",
              "content": "<p>How did you implement the lag? The drop is quite crazy. I suspect some wrong doing here. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3032424,
                  "author_name": "dinezra11",
                  "author_url": "",
                  "post_date": "10/30/2024 21:45:02",
                  "content": "<p><a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a> I created the lags feature with a simple shift:<br>\npl.col('responder_6').shift(1).over(partition_by='symbol_id',order_by=['date_id', 'time_id']).alias('lag_1time responder_6')</p>\n<p>But I suspect you are right and I did something wrong in using the API. I got a feeling that I didn't understand the API completely. Only today I realized that the actual test data is served from the API in batches and not all in one go, so I think the loading part of the lag dataframe gone wrong.<br>\nDo we know if it is guaranteed that each batch from API has records from the same date_id and time_id? And it is guaranteed that the batches are served chronologically?</p>\n<p>I'll try to debug my code a bit more to understand it better.<br>\nAppreciate your help!</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3032468,
                      "author_name": "shiyili",
                      "author_url": "",
                      "post_date": "10/30/2024 23:27:20",
                      "content": "<p>Your shit(1).over() seems to be wrong. You don't lag the responders by 1 time-step, instead you should lag them by 1 day.</p>\n<p>If you do the following plot, you will find the raw responder_6 and the lagged are not matching.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F3dd6c0a2b7c65d7d8ca2ed78aa894621%2FScreenshot%202024-10-31%20at%2000.25.19.png?generation=1730330724207812&amp;alt=media\" alt=\"\"></p>\n<p>You should instead do:</p>\n<pre><code>df_all = pl.read_parquet(train_partitions, =[, , , ]).with_columns(\n        pl.col().shift().([, ]).()\n    )\n</code></pre>\n<p>which will generate matching series:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F54d4464f5be7cabaef5d0c6866d1a395%2FScreenshot%202024-10-31%20at%2000.26.39.png?generation=1730330808063141&amp;alt=media\" alt=\"\"></p>\n<p>Plotting code: </p>\n<pre><code>data_dir = Path(\"./data\")\ntrain_partitions = [data_dir / f\"train.parquet/partition_id={i}/part-0.parquet\"  i  range(, )]\n\ndf_all = pl.read_parquet(train_partitions, =[, , , ]).with_columns(\n        pl.col().shift().(partition_by=,order_by=[, ]).()\n    )\n\np1 = df_all.((pl.col()==)&amp;(pl.col()==)).([, ]).to_pandas()\np2 = df_all.((pl.col()==)&amp;(pl.col()==)).([, ]).to_pandas()\n\nplt.figure(figsize=(, ))\nplt.plot(p1[], p1[], label=)\nplt.plot(p2[], p2[], label=)\n\nplt.legend()\nplt.()\n</code></pre>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3032495,
                          "author_name": "dinezra11",
                          "author_url": "",
                          "post_date": "10/31/2024 00:49:34",
                          "content": "<p>Oh wow, seems like a severe mistake. Thank you for pointing it out!<br>\nCould you explain me the reason why we don't lag by time step, and why it's better to lag by date?<br>\nI want to have a deep understanding in time-series best practices.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3032504,
                              "author_name": "jackvd",
                              "author_url": "",
                              "post_date": "10/31/2024 01:37:05",
                              "content": "<p>It’s not better to lag by date, it’s just how lags is provided in the evaluation api. </p>",
                              "votes": null,
                              "replies": []
                            },
                            {
                              "id": 3032703,
                              "author_name": "shiyili",
                              "author_url": "",
                              "post_date": "10/31/2024 08:52:11",
                              "content": "<p>It’s not a best practice. You do so to bc this is how the api serves data </p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3032944,
                                  "author_name": "dinezra11",
                                  "author_url": "",
                                  "post_date": "10/31/2024 15:16:04",
                                  "content": "<p>I see. Thank you all very much for the help and the clarifications! I feel like I learnt a lot now.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 3032264,
              "author_name": "chen2022gz",
              "author_url": "",
              "post_date": "10/30/2024 18:15:42",
              "content": "<p>i also face the same problem</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3081039,
          "author_name": "akomwins",
          "author_url": "",
          "post_date": "12/26/2024 06:15:10",
          "content": "<p>I added extra feature_00_lags to features_78_lags besides the original features, but it seems it doesn work, have you tried it?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3084180,
              "author_name": "goldenlock",
              "author_url": "",
              "post_date": "12/30/2024 13:56:48",
              "content": "<p>If I made it correctly I tested with 4 feature lags using tree models, it not change much offline and got much worse LB.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3032013,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "10/30/2024 13:00:59",
      "content": "<p>You can check out this notebook for a solid example on using <code>lag_1</code> on responders in inference. </p>\n<p><a href=\"https://www.kaggle.com/code/motono0223/js24-inference-gbdt-with-lags-singlemodel\" target=\"_blank\">https://www.kaggle.com/code/motono0223/js24-inference-gbdt-with-lags-singlemodel</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3032428,
          "author_name": "dinezra11",
          "author_url": "",
          "post_date": "10/30/2024 21:48:00",
          "content": "<p>Thank you very much! I'll check it out. I think I misunderstood the API part and I believe this notebook will clear up my confusion</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3031421": "I was wondering if anyone made lag features that helped his LB score improve. So far I tried do some feature engineering and create lag features but it always led to far worse score. I wonder if lags are useful or I should avoid them in this competition.\n\nThanks in advance!",
    "3031437": "Many features are lagged (or \"semi-lagged\") from each other, so lagging all features will not be helpful (only cause redundancy). Some features are not suitable for making it lagged. You need to properly choose the right feature to lag the right timesteps.",
    "3031453": "Yes, so far I tried to lag only the target column (responder_6), but it didn't lead to good results (dropped my 0.43 LB score significantly to -0.024!), which was strange to me because to my logic and understanding the previous values of the variable we are trying to predict are usually very important in time-series problems.",
    "3031468": "this drop in LB score is insane!! you must have introduced a bug to your training  or predicting code while doing feature engineering",
    "3031494": "How did you implement the lag? The drop is quite crazy. I suspect some wrong doing here.",
    "3032013": "You can check out this notebook for a solid example on using `lag_1` on responders in inference. \n\nhttps://www.kaggle.com/code/motono0223/js24-inference-gbdt-with-lags-singlemodel",
    "3032264": "i also face the same problem",
    "3032424": "shiyili I created the lags feature with a simple shift:\npl.col('responder_6').shift(1).over(partition_by='symbol_id',order_by=['date_id', 'time_id']).alias('lag_1time responder_6')\n\nBut I suspect you are right and I did something wrong in using the API. I got a feeling that I didn't understand the API completely. Only today I realized that the actual test data is served from the API in batches and not all in one go, so I think the loading part of the lag dataframe gone wrong.\nDo we know if it is guaranteed that each batch from API has records from the same date_id and time_id? And it is guaranteed that the batches are served chronologically?\n\nI'll try to debug my code a bit more to understand it better.\nAppreciate your help!",
    "3032428": "Thank you very much! I'll check it out. I think I misunderstood the API part and I believe this notebook will clear up my confusion",
    "3032468": "Your shit(1).over() seems to be wrong. You don't lag the responders by 1 time-step, instead you should lag them by 1 day.\n\nIf you do the following plot, you will find the raw responder_6 and the lagged are not matching.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F3dd6c0a2b7c65d7d8ca2ed78aa894621%2FScreenshot%202024-10-31%20at%2000.25.19.png?generation=1730330724207812&alt=media)\n\nYou should instead do:\n\n```\ndf_all = pl.read_parquet(train_partitions, columns=['date_id', 'time_id', 'symbol_id', 'responder_6']).with_columns(\n        pl.col('responder_6').shift(1).over(['symbol_id', 'time_id']).alias('lag_1time responder_6')\n    )\n```\n\nwhich will generate matching series:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F54d4464f5be7cabaef5d0c6866d1a395%2FScreenshot%202024-10-31%20at%2000.26.39.png?generation=1730330808063141&alt=media)\n\nPlotting code: \n\n```\ndata_dir = Path(\"./data\")\ntrain_partitions = [data_dir / f\"train.parquet/partition_id={i}/part-0.parquet\" for i in range(9, 10)]\n\ndf_all = pl.read_parquet(train_partitions, columns=['date_id', 'time_id', 'symbol_id', 'responder_6']).with_columns(\n        pl.col('responder_6').shift(1).over(partition_by='symbol_id',order_by=['date_id', 'time_id']).alias('lag_1time responder_6')\n    )\n\np1 = df_all.filter((pl.col('date_id')==1530)&(pl.col('symbol_id')==0)).select(['time_id', 'responder_6']).to_pandas()\np2 = df_all.filter((pl.col('date_id')==1531)&(pl.col('symbol_id')==0)).select(['time_id', 'lag_1time responder_6']).to_pandas()\n\nplt.figure(figsize=(12, 3))\nplt.plot(p1['time_id'], p1['responder_6'], label='responder_6')\nplt.plot(p2['time_id'], p2['lag_1time responder_6'], label='lag_1time responder_6')\n\nplt.legend()\nplt.show()\n```",
    "3032495": "Oh wow, seems like a severe mistake. Thank you for pointing it out!\nCould you explain me the reason why we don't lag by time step, and why it's better to lag by date?\nI want to have a deep understanding in time-series best practices.",
    "3032504": "It’s not better to lag by date, it’s just how lags is provided in the evaluation api.",
    "3032703": "It’s not a best practice. You do so to bc this is how the api serves data",
    "3032944": "I see. Thank you all very much for the help and the clarifications! I feel like I learnt a lot now.",
    "3081039": "I added extra feature_00_lags to features_78_lags besides the original features, but it seems it doesn work, have you tried it?",
    "3084180": "If I made it correctly I tested with 4 feature lags using tree models, it not change much offline and got much worse LB."
  },
  "source": "meta"
}