{
  "id": 205708,
  "title": "multiclass format is not supported",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205708",
  "author_name": "SHYJohn",
  "post_date": "2020-12-21T13:44:53.756000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am using LighGBM to train my model, in addition to that I am using hyperparameter optimisation using <code>RandomizedSearchCV</code> from <code>sklearn.model_selection</code>. When I am trying to fit my model, I encounter the following error: </p>\n<pre><code>ValueError: multiclass format is not supported\n</code></pre>\n<p>I wonder what has happened? When I looked into Google, it says I should not use <code>roc</code> as metric. But I guess this competition needs them?</p>\n<p>Here is my MWE: </p>\n<pre><code>import riiideducation\n\nimport numpy as np \nimport pandas as pd \n\nimport lightgbm as lgb\n\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import RandomizedSearchCV, GridSearchCV\nfrom scipy.stats import randint as sp_randint\nfrom scipy.stats import uniform as sp_uniform\n\n'''\nOther code re: cleaning data here. \n'''\nlgb_train = lgb.Dataset(X_train_df, y_train_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\nlgb_eval = lgb.Dataset(X_val_df, y_val_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\n\n# param values c.f. https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2\nfit_param = {\"eval_metric\" : 'auc'}\n\nparam = {'num_leaves': sp_randint(30, 400), 'max_bin':sp_randint(300, 800), 'min_child_weight': [1e-5, 1e-3, 1e-2, 1e-1, 1, 1e1, 1e2, 1e3, 1e4], \n         'feature_fraction': sp_uniform(0, 1), 'bagging_fraction': sp_uniform(0, 1), \n         'objective': ['binary'], 'max_depth': [-1], \n         'learning_rate': [0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09], \"boosting_type\": [\"gbdt\"], \"bagging_seed\": sp_randint(5, 30), \n         'eval_metric': ['auc'], \"verbosity\": [-1], \n         'reg_alpha': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], 'reg_lambda': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], \n         'random_state': [47]}\n\nm1 = lgb.LGBMRegressor(valid_sets = [lgb_train, lgb_eval], verbose_eval = 30, num_boost_round = 10000, early_stopping_rounds = 30, n_jobs=4, n_estimators=3000)\n'''\nHyperparameter optimisation\n'''\n# Code from https://www.kaggle.com/rtatman/lightgbm-hyperparameter-optimisation-lb-0-761#Model-fitting-with-HyperParameter-optimisation\n#This parameter defines the number of HP points to be tested\nn_HP_points_to_test = 100\n\ngsLGBM = RandomizedSearchCV(\n    estimator=m1, param_distributions=param, \n    n_iter=n_HP_points_to_test,\n    scoring='roc_auc',\n    cv=3,\n    refit=True,\n    random_state=47,\n    verbose=True)\n\ngsLGBM.fit(X_train_df, y_train_df, eval_set = (X_val_df, y_val_df), eval_metric = 'multi:softmax')\n</code></pre>",
  "messages": [
    {
      "id": 1121257,
      "postDate": "2020-12-21T13:44:53.757Z",
      "content": "<p>I am using LighGBM to train my model, in addition to that I am using hyperparameter optimisation using <code>RandomizedSearchCV</code> from <code>sklearn.model_selection</code>. When I am trying to fit my model, I encounter the following error: </p>\n<pre><code>ValueError: multiclass format is not supported\n</code></pre>\n<p>I wonder what has happened? When I looked into Google, it says I should not use <code>roc</code> as metric. But I guess this competition needs them?</p>\n<p>Here is my MWE: </p>\n<pre><code>import riiideducation\n\nimport numpy as np \nimport pandas as pd \n\nimport lightgbm as lgb\n\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import RandomizedSearchCV, GridSearchCV\nfrom scipy.stats import randint as sp_randint\nfrom scipy.stats import uniform as sp_uniform\n\n'''\nOther code re: cleaning data here. \n'''\nlgb_train = lgb.Dataset(X_train_df, y_train_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\nlgb_eval = lgb.Dataset(X_val_df, y_val_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\n\n# param values c.f. https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2\nfit_param = {\"eval_metric\" : 'auc'}\n\nparam = {'num_leaves': sp_randint(30, 400), 'max_bin':sp_randint(300, 800), 'min_child_weight': [1e-5, 1e-3, 1e-2, 1e-1, 1, 1e1, 1e2, 1e3, 1e4], \n         'feature_fraction': sp_uniform(0, 1), 'bagging_fraction': sp_uniform(0, 1), \n         'objective': ['binary'], 'max_depth': [-1], \n         'learning_rate': [0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09], \"boosting_type\": [\"gbdt\"], \"bagging_seed\": sp_randint(5, 30), \n         'eval_metric': ['auc'], \"verbosity\": [-1], \n         'reg_alpha': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], 'reg_lambda': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], \n         'random_state': [47]}\n\nm1 = lgb.LGBMRegressor(valid_sets = [lgb_train, lgb_eval], verbose_eval = 30, num_boost_round = 10000, early_stopping_rounds = 30, n_jobs=4, n_estimators=3000)\n'''\nHyperparameter optimisation\n'''\n# Code from https://www.kaggle.com/rtatman/lightgbm-hyperparameter-optimisation-lb-0-761#Model-fitting-with-HyperParameter-optimisation\n#This parameter defines the number of HP points to be tested\nn_HP_points_to_test = 100\n\ngsLGBM = RandomizedSearchCV(\n    estimator=m1, param_distributions=param, \n    n_iter=n_HP_points_to_test,\n    scoring='roc_auc',\n    cv=3,\n    refit=True,\n    random_state=47,\n    verbose=True)\n\ngsLGBM.fit(X_train_df, y_train_df, eval_set = (X_val_df, y_val_df), eval_metric = 'multi:softmax')\n</code></pre>",
      "rawMarkdown": "I am using LighGBM to train my model, in addition to that I am using hyperparameter optimisation using `RandomizedSearchCV` from `sklearn.model_selection`. When I am trying to fit my model, I encounter the following error: \n```\nValueError: multiclass format is not supported\n```\nI wonder what has happened? When I looked into Google, it says I should not use `roc` as metric. But I guess this competition needs them?\n\nHere is my MWE: \n```\nimport riiideducation\n\nimport numpy as np \nimport pandas as pd \n\nimport lightgbm as lgb\n\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import RandomizedSearchCV, GridSearchCV\nfrom scipy.stats import randint as sp_randint\nfrom scipy.stats import uniform as sp_uniform\n\n'''\nOther code re: cleaning data here. \n'''\nlgb_train = lgb.Dataset(X_train_df, y_train_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\nlgb_eval = lgb.Dataset(X_val_df, y_val_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\n\n# param values c.f. https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2\nfit_param = {\"eval_metric\" : 'auc'}\n\nparam = {'num_leaves': sp_randint(30, 400), 'max_bin':sp_randint(300, 800), 'min_child_weight': [1e-5, 1e-3, 1e-2, 1e-1, 1, 1e1, 1e2, 1e3, 1e4], \n         'feature_fraction': sp_uniform(0, 1), 'bagging_fraction': sp_uniform(0, 1), \n         'objective': ['binary'], 'max_depth': [-1], \n         'learning_rate': [0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09], \"boosting_type\": [\"gbdt\"], \"bagging_seed\": sp_randint(5, 30), \n         'eval_metric': ['auc'], \"verbosity\": [-1], \n         'reg_alpha': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], 'reg_lambda': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], \n         'random_state': [47]}\n\nm1 = lgb.LGBMRegressor(valid_sets = [lgb_train, lgb_eval], verbose_eval = 30, num_boost_round = 10000, early_stopping_rounds = 30, n_jobs=4, n_estimators=3000)\n'''\nHyperparameter optimisation\n'''\n# Code from https://www.kaggle.com/rtatman/lightgbm-hyperparameter-optimisation-lb-0-761#Model-fitting-with-HyperParameter-optimisation\n#This parameter defines the number of HP points to be tested\nn_HP_points_to_test = 100\n\ngsLGBM = RandomizedSearchCV(\n    estimator=m1, param_distributions=param, \n    n_iter=n_HP_points_to_test,\n    scoring='roc_auc',\n    cv=3,\n    refit=True,\n    random_state=47,\n    verbose=True)\n\ngsLGBM.fit(X_train_df, y_train_df, eval_set = (X_val_df, y_val_df), eval_metric = 'multi:softmax')\n```"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1121257": "I am using LighGBM to train my model, in addition to that I am using hyperparameter optimisation using `RandomizedSearchCV` from `sklearn.model_selection`. When I am trying to fit my model, I encounter the following error: \n```\nValueError: multiclass format is not supported\n```\nI wonder what has happened? When I looked into Google, it says I should not use `roc` as metric. But I guess this competition needs them?\n\nHere is my MWE: \n```\nimport riiideducation\n\nimport numpy as np \nimport pandas as pd \n\nimport lightgbm as lgb\n\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import RandomizedSearchCV, GridSearchCV\nfrom scipy.stats import randint as sp_randint\nfrom scipy.stats import uniform as sp_uniform\n\n'''\nOther code re: cleaning data here. \n'''\nlgb_train = lgb.Dataset(X_train_df, y_train_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\nlgb_eval = lgb.Dataset(X_val_df, y_val_df, categorical_feature = ['prior_question_had_explanation'], free_raw_data=False)\n\n# param values c.f. https://www.kaggle.com/zephyrwang666/riiid-lgbm-bagging2\nfit_param = {\"eval_metric\" : 'auc'}\n\nparam = {'num_leaves': sp_randint(30, 400), 'max_bin':sp_randint(300, 800), 'min_child_weight': [1e-5, 1e-3, 1e-2, 1e-1, 1, 1e1, 1e2, 1e3, 1e4], \n         'feature_fraction': sp_uniform(0, 1), 'bagging_fraction': sp_uniform(0, 1), \n         'objective': ['binary'], 'max_depth': [-1], \n         'learning_rate': [0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09], \"boosting_type\": [\"gbdt\"], \"bagging_seed\": sp_randint(5, 30), \n         'eval_metric': ['auc'], \"verbosity\": [-1], \n         'reg_alpha': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], 'reg_lambda': [0, 1e-1, 1, 2, 5, 7, 10, 50, 100], \n         'random_state': [47]}\n\nm1 = lgb.LGBMRegressor(valid_sets = [lgb_train, lgb_eval], verbose_eval = 30, num_boost_round = 10000, early_stopping_rounds = 30, n_jobs=4, n_estimators=3000)\n'''\nHyperparameter optimisation\n'''\n# Code from https://www.kaggle.com/rtatman/lightgbm-hyperparameter-optimisation-lb-0-761#Model-fitting-with-HyperParameter-optimisation\n#This parameter defines the number of HP points to be tested\nn_HP_points_to_test = 100\n\ngsLGBM = RandomizedSearchCV(\n    estimator=m1, param_distributions=param, \n    n_iter=n_HP_points_to_test,\n    scoring='roc_auc',\n    cv=3,\n    refit=True,\n    random_state=47,\n    verbose=True)\n\ngsLGBM.fit(X_train_df, y_train_df, eval_set = (X_val_df, y_val_df), eval_metric = 'multi:softmax')\n```"
  }
}