{
  "id": 540472,
  "title": "Dealing with clunky logs in boosted tree models with custom eval metrics ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/540472",
  "author_name": "Ravi Ramakrishnan",
  "post_date": "2024-10-14T17:43:51.984000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>This competition uses a weighted r2score metric that could be considered as a custom metric for model CV scores. </p>\n<p>Boosted tree models like lightgbm and xgboost often result in clunky log messages while working with custom objectives and eval metrics. The latest version of lightgbm (4.5.0) resolves this issue but is not directly present in the Kaggle base environment as on date. </p>\n<p>One may resort to the below code snippet to control log messages while using the present version of lightgbm4.2.0 - </p>\n<pre><code> logging, lightgbm  lgb\n\n\n :\n    \n\n     ():\n        self.logger = logging.getLogger(logging_lbl)\n        self.logger.setLevel(logging.ERROR)\n\n     ():\n        \n\n     ():\n        \n\n     ():\n        self.logger.error(message)\n\nl = MyLogger()\nl.init(logging_lbl = )\nlgb.register_logger(l)\n</code></pre>\n<p>One could also control logging in xgboost as below-</p>\n<pre><code>\n xgboost  xgb, logging \n\n handler  logging.root.handlers[:]:\n    logging.root.removeHandler(handler)\n\nlogger = logging.getLogger(__name__)\nlogger.setLevel(logging.ERROR)\nformatter = logging.Formatter()\n\nstdout_handler = logging.StreamHandler(sys.stdout)\nstdout_handler.setLevel(logging.INFO)\nstdout_handler.setFormatter(formatter)\n\nfile_handler = logging.FileHandler()\nfile_handler.setLevel(logging.ERROR)\nfile_handler.setFormatter(formatter)\n\nlogger.addHandler(file_handler)\nlogger.addHandler(stdout_handler)\n\n (xgb.callback.TrainingCallback):\n    \n\n     ():\n        self.epoch_log_interval = epoch_log_interval\n\n     ():\n\n         self.epoch_log_interval &lt;= :\n            \n\n         (epoch %  self.epoch_log_interval == ):\n             data, metric  evals_log.items():\n                 metric_name, log  metric.items():\n                    score = log[-][]  (log[-], )  log[-]\n                    logger.info()\n\n         \n</code></pre>\n<p>Please feel free to use the latest versions of lightgbm, scikit-learn and xgboost in your pipelines from here-<br>\n<a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1/notebook\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1/notebook</a></p>\n<p>The above code snippet and the latest versions of lightgbm, scikit-learn and xgboost packages are present for ready-usage</p>\n<p>All the best!</p>",
  "messages": [
    {
      "id": 3017271,
      "postDate": "2024-10-14T17:43:51.983Z",
      "content": "<p>Hello all,</p>\n<p>This competition uses a weighted r2score metric that could be considered as a custom metric for model CV scores. </p>\n<p>Boosted tree models like lightgbm and xgboost often result in clunky log messages while working with custom objectives and eval metrics. The latest version of lightgbm (4.5.0) resolves this issue but is not directly present in the Kaggle base environment as on date. </p>\n<p>One may resort to the below code snippet to control log messages while using the present version of lightgbm4.2.0 - </p>\n<pre><code> logging, lightgbm  lgb\n\n\n :\n    \n\n     ():\n        self.logger = logging.getLogger(logging_lbl)\n        self.logger.setLevel(logging.ERROR)\n\n     ():\n        \n\n     ():\n        \n\n     ():\n        self.logger.error(message)\n\nl = MyLogger()\nl.init(logging_lbl = )\nlgb.register_logger(l)\n</code></pre>\n<p>One could also control logging in xgboost as below-</p>\n<pre><code>\n xgboost  xgb, logging \n\n handler  logging.root.handlers[:]:\n    logging.root.removeHandler(handler)\n\nlogger = logging.getLogger(__name__)\nlogger.setLevel(logging.ERROR)\nformatter = logging.Formatter()\n\nstdout_handler = logging.StreamHandler(sys.stdout)\nstdout_handler.setLevel(logging.INFO)\nstdout_handler.setFormatter(formatter)\n\nfile_handler = logging.FileHandler()\nfile_handler.setLevel(logging.ERROR)\nfile_handler.setFormatter(formatter)\n\nlogger.addHandler(file_handler)\nlogger.addHandler(stdout_handler)\n\n (xgb.callback.TrainingCallback):\n    \n\n     ():\n        self.epoch_log_interval = epoch_log_interval\n\n     ():\n\n         self.epoch_log_interval &lt;= :\n            \n\n         (epoch %  self.epoch_log_interval == ):\n             data, metric  evals_log.items():\n                 metric_name, log  metric.items():\n                    score = log[-][]  (log[-], )  log[-]\n                    logger.info()\n\n         \n</code></pre>\n<p>Please feel free to use the latest versions of lightgbm, scikit-learn and xgboost in your pipelines from here-<br>\n<a href=\"https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1/notebook\" target=\"_blank\">https://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1/notebook</a></p>\n<p>The above code snippet and the latest versions of lightgbm, scikit-learn and xgboost packages are present for ready-usage</p>\n<p>All the best!</p>",
      "rawMarkdown": "Hello all,\n\nThis competition uses a weighted r2score metric that could be considered as a custom metric for model CV scores. \n\nBoosted tree models like lightgbm and xgboost often result in clunky log messages while working with custom objectives and eval metrics. The latest version of lightgbm (4.5.0) resolves this issue but is not directly present in the Kaggle base environment as on date. \n\nOne may resort to the below code snippet to control log messages while using the present version of lightgbm4.2.0 - \n\n```python\nimport logging, lightgbm as lgb\n\n# Customizing logging for LGBM\nclass MyLogger:\n    \"\"\"\n    This class helps to suppress logs in lightgbm \n    Source - https://github.com/microsoft/LightGBM/issues/6014\n    \"\"\"\n\n    def init(self, logging_lbl: str):\n        self.logger = logging.getLogger(logging_lbl)\n        self.logger.setLevel(logging.ERROR)\n\n    def info(self, message):\n        pass\n\n    def warning(self, message):\n        pass\n\n    def error(self, message):\n        self.logger.error(message)\n\nl = MyLogger()\nl.init(logging_lbl = \"lightgbm_custom\")\nlgb.register_logger(l)\n```\n\nOne could also control logging in xgboost as below-\n\n```python\n# Customizing logging for XGBoost\nimport xgboost as xgb, logging \n\nfor handler in logging.root.handlers[:]:\n    logging.root.removeHandler(handler)\n\nlogger = logging.getLogger(__name__)\nlogger.setLevel(logging.ERROR)\nformatter = logging.Formatter('%(asctime)s | %(levelname)s | %(message)s')\n\nstdout_handler = logging.StreamHandler(sys.stdout)\nstdout_handler.setLevel(logging.INFO)\nstdout_handler.setFormatter(formatter)\n\nfile_handler = logging.FileHandler(f'xgb_optimize.log')\nfile_handler.setLevel(logging.ERROR)\nfile_handler.setFormatter(formatter)\n\nlogger.addHandler(file_handler)\nlogger.addHandler(stdout_handler)\n\nclass XGBLogging(xgb.callback.TrainingCallback):\n    \"\"\"log train logs to file\"\"\"\n\n    def __init__(self, epoch_log_interval=100):\n        self.epoch_log_interval = epoch_log_interval\n\n    def after_iteration(self, model, epoch:int,\n                        evals_log:xgb.callback.TrainingCallback.EvalsLog\n                        ):\n\n        if self.epoch_log_interval <= 0:\n            pass\n\n        elif (epoch %  self.epoch_log_interval == 0):\n            for data, metric in evals_log.items():\n                for metric_name, log in metric.items():\n                    score = log[-1][0] if isinstance(log[-1], tuple) else log[-1]\n                    logger.info(f\"XGBLogging epoch {epoch} dataset {data} {metric_name} {score}\")\n\n        return False\n```\n\nPlease feel free to use the latest versions of lightgbm, scikit-learn and xgboost in your pipelines from here-\nhttps://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1/notebook\n\nThe above code snippet and the latest versions of lightgbm, scikit-learn and xgboost packages are present for ready-usage\n\nAll the best!\n",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3017271": "Hello all,\n\nThis competition uses a weighted r2score metric that could be considered as a custom metric for model CV scores. \n\nBoosted tree models like lightgbm and xgboost often result in clunky log messages while working with custom objectives and eval metrics. The latest version of lightgbm (4.5.0) resolves this issue but is not directly present in the Kaggle base environment as on date. \n\nOne may resort to the below code snippet to control log messages while using the present version of lightgbm4.2.0 - \n\n```python\nimport logging, lightgbm as lgb\n\n# Customizing logging for LGBM\nclass MyLogger:\n    \"\"\"\n    This class helps to suppress logs in lightgbm \n    Source - https://github.com/microsoft/LightGBM/issues/6014\n    \"\"\"\n\n    def init(self, logging_lbl: str):\n        self.logger = logging.getLogger(logging_lbl)\n        self.logger.setLevel(logging.ERROR)\n\n    def info(self, message):\n        pass\n\n    def warning(self, message):\n        pass\n\n    def error(self, message):\n        self.logger.error(message)\n\nl = MyLogger()\nl.init(logging_lbl = \"lightgbm_custom\")\nlgb.register_logger(l)\n```\n\nOne could also control logging in xgboost as below-\n\n```python\n# Customizing logging for XGBoost\nimport xgboost as xgb, logging \n\nfor handler in logging.root.handlers[:]:\n    logging.root.removeHandler(handler)\n\nlogger = logging.getLogger(__name__)\nlogger.setLevel(logging.ERROR)\nformatter = logging.Formatter('%(asctime)s | %(levelname)s | %(message)s')\n\nstdout_handler = logging.StreamHandler(sys.stdout)\nstdout_handler.setLevel(logging.INFO)\nstdout_handler.setFormatter(formatter)\n\nfile_handler = logging.FileHandler(f'xgb_optimize.log')\nfile_handler.setLevel(logging.ERROR)\nfile_handler.setFormatter(formatter)\n\nlogger.addHandler(file_handler)\nlogger.addHandler(stdout_handler)\n\nclass XGBLogging(xgb.callback.TrainingCallback):\n    \"\"\"log train logs to file\"\"\"\n\n    def __init__(self, epoch_log_interval=100):\n        self.epoch_log_interval = epoch_log_interval\n\n    def after_iteration(self, model, epoch:int,\n                        evals_log:xgb.callback.TrainingCallback.EvalsLog\n                        ):\n\n        if self.epoch_log_interval <= 0:\n            pass\n\n        elif (epoch %  self.epoch_log_interval == 0):\n            for data, metric in evals_log.items():\n                for metric_name, log in metric.items():\n                    score = log[-1][0] if isinstance(log[-1], tuple) else log[-1]\n                    logger.info(f\"XGBLogging epoch {epoch} dataset {data} {metric_name} {score}\")\n\n        return False\n```\n\nPlease feel free to use the latest versions of lightgbm, scikit-learn and xgboost in your pipelines from here-\nhttps://www.kaggle.com/code/ravi20076/janestreet2024-imports-v1/notebook\n\nThe above code snippet and the latest versions of lightgbm, scikit-learn and xgboost packages are present for ready-usage\n\nAll the best!\n"
  }
}