{
  "id": 420616,
  "title": "Efficiency track 18th place (25th public): length of test dataframe, groupby, mean",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420616",
  "author_name": "",
  "post_date": "2023-07-01T16:37:41.570714100Z",
  "votes": 8,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I wanted to make simple baselines. The first baseline is using the average of correct answers for each question and finding a threshold, scoring .659. For my second baseline, I wanted to incorporate some more information. Considering the efficiency track of the competition, I thought \"what information can I get and use efficiently, ideally without even iterating through all the rows of data?\". My answer was that I can get the length of the test dataframe. I wanted to do a simple groupby… mean, but len(df) is not discrete enough. with chatgpt I found KBinsDiscretizer from sklearn and used that to bin len(df) feature so I could use groupby. </p>\n<p>Here is the code, scoring .671 public .675 private, currently 25th on efficiency leaderboard: </p>\n<pre><code> MeanKBD:\n    def __init__(self): \n        pass\n\n    def fit(self, trn, trn_labels, raw=None):\n         raw is not None: \n            raw = Path(raw)\n            trn = pd.read_csv(raw/)\n            trn_labels = pd.read_csv(raw/)\n        trn_labels = add_basic_cols_to_ss(trn_labels)\n        tl = trn_labels.set_index([, ])\n        trn[] = trn.level_group.map({: , : , : })\n        tl[] = trn.groupby([, ])[].count()\n\n        ##### Discretize length  session ########\n        kbd_dict = {}\n        for level  [, , ]: \n            tll = tl[tl.index.get_level_values() == level]\n            tll.row_count.values.reshape(, ).shape\n            kbd = KBinsDiscretizer(n_bins=)\n            tll = tll.assign(kbd=kbd.fit_transform(tll.row_count.values.reshape(, )).argmax(axis=).astype(int))\n            tl.loc[tl.index.get_level_values() == level, ] = tll[]\n            kbd_dict[level] = kbd\n        tl[] = tl[].astype(int)\n        self.kbd_dict = kbd_dict\n        self.q_kbd_dict = tl.groupby([, ])[].mean().to_dict()\n\n        ##### Find best threshold #####\n        prob = tl.apply(lambda row: (row[], row[]), axis=).map(self.q_kbd_dict)\n        best_score = \n        thresholds, scores = [], []\n        for th  np.arange(, , ): \n            score = f1_score(tl[], (prob &gt; th).astype(int), average=)\n            thresholds.append(th)\n            scores.append(score)\n             score &gt; best_score: \n                best_score = score\n                best_th = th\n        print(f)\n        # sns.lineplot(x=thresholds, y=scores, marker=)\n        # plt.show()\n        self.best_th = best_th\n\n    def predict(self, env, fast=): \n         fast: \n            tl = env.test_labels.set_index([, ])\n            trn = env.test\n            tl[] = trn.groupby([, ])[].count()\n            for level  [, , ]:\n                mask = tl.index.get_level_values() == level\n                kbd = self.kbd_dict[level].transform(tl.loc[mask, ].values.reshape(, ))\n                tl.loc[mask, ] = kbd.argmax(axis=)\n            prob = tl.apply(lambda row: (row[], row[]), axis=).map(self.q_kbd_dict)\n            tl[] = (prob &gt; self.best_th).astype(int)\n            env.ss = tl\n            return None\n\n        for i, (test, ss)  enumerate(env.iter_test()): \n            tmp = ss.copy()\n            tmp[] = len(test)\n            tmp[] = tmp.session_id.str.split().str[].str[:].astype(int)\n            tmp[] = self.kbd_dict[i % ].transform(tmp[].values.reshape(, )).argmax(axis=)\n            prob = tmp.apply(lambda row: (row[], row[]), axis=).map(self.q_kbd_dict)\n\n            ss[] = (prob &gt; self.best_th).astype(int)\n            env.predict(ss)\n\n joblib\n gameplay  MeanKBD\n\nm = MeanKBD()\nm.fit(None, None, raw=)\njoblib.dump(m, )\n\n###################### Put the rest  the   an inference notebook ################\n joblib\n gameplay  MeanKBD\n\nm = joblib.load()\n jo_wilder_310\nenv = jo_wilder_310.make_env()\nm.predict(env)\n</code></pre>",
  "messages": [
    {
      "id": "2325864",
      "postDate": "07/01/2023 16:37:41",
      "content": "<p>I wanted to make simple baselines. The first baseline is using the average of correct answers for each question and finding a threshold, scoring .659. For my second baseline, I wanted to incorporate some more information. Considering the efficiency track of the competition, I thought \"what information can I get and use efficiently, ideally without even iterating through all the rows of data?\". My answer was that I can get the length of the test dataframe. I wanted to do a simple groupby… mean, but len(df) is not discrete enough. with chatgpt I found KBinsDiscretizer from sklearn and used that to bin len(df) feature so I could use groupby. </p>\n<p>Here is the code, scoring .671 public .675 private, currently 25th on efficiency leaderboard: </p>\n<pre><code> MeanKBD:\n    def __init__(self): \n        pass\n\n    def fit(self, trn, trn_labels, raw=None):\n         raw is not None: \n            raw = Path(raw)\n            trn = pd.read_csv(raw/)\n            trn_labels = pd.read_csv(raw/)\n        trn_labels = add_basic_cols_to_ss(trn_labels)\n        tl = trn_labels.set_index([, ])\n        trn[] = trn.level_group.map({: , : , : })\n        tl[] = trn.groupby([, ])[].count()\n\n        ##### Discretize length  session ########\n        kbd_dict = {}\n        for level  [, , ]: \n            tll = tl[tl.index.get_level_values() == level]\n            tll.row_count.values.reshape(, ).shape\n            kbd = KBinsDiscretizer(n_bins=)\n            tll = tll.assign(kbd=kbd.fit_transform(tll.row_count.values.reshape(, )).argmax(axis=).astype(int))\n            tl.loc[tl.index.get_level_values() == level, ] = tll[]\n            kbd_dict[level] = kbd\n        tl[] = tl[].astype(int)\n        self.kbd_dict = kbd_dict\n        self.q_kbd_dict = tl.groupby([, ])[].mean().to_dict()\n\n        ##### Find best threshold #####\n        prob = tl.apply(lambda row: (row[], row[]), axis=).map(self.q_kbd_dict)\n        best_score = \n        thresholds, scores = [], []\n        for th  np.arange(, , ): \n            score = f1_score(tl[], (prob &gt; th).astype(int), average=)\n            thresholds.append(th)\n            scores.append(score)\n             score &gt; best_score: \n                best_score = score\n                best_th = th\n        print(f)\n        # sns.lineplot(x=thresholds, y=scores, marker=)\n        # plt.show()\n        self.best_th = best_th\n\n    def predict(self, env, fast=): \n         fast: \n            tl = env.test_labels.set_index([, ])\n            trn = env.test\n            tl[] = trn.groupby([, ])[].count()\n            for level  [, , ]:\n                mask = tl.index.get_level_values() == level\n                kbd = self.kbd_dict[level].transform(tl.loc[mask, ].values.reshape(, ))\n                tl.loc[mask, ] = kbd.argmax(axis=)\n            prob = tl.apply(lambda row: (row[], row[]), axis=).map(self.q_kbd_dict)\n            tl[] = (prob &gt; self.best_th).astype(int)\n            env.ss = tl\n            return None\n\n        for i, (test, ss)  enumerate(env.iter_test()): \n            tmp = ss.copy()\n            tmp[] = len(test)\n            tmp[] = tmp.session_id.str.split().str[].str[:].astype(int)\n            tmp[] = self.kbd_dict[i % ].transform(tmp[].values.reshape(, )).argmax(axis=)\n            prob = tmp.apply(lambda row: (row[], row[]), axis=).map(self.q_kbd_dict)\n\n            ss[] = (prob &gt; self.best_th).astype(int)\n            env.predict(ss)\n\n joblib\n gameplay  MeanKBD\n\nm = MeanKBD()\nm.fit(None, None, raw=)\njoblib.dump(m, )\n\n###################### Put the rest  the   an inference notebook ################\n joblib\n gameplay  MeanKBD\n\nm = joblib.load()\n jo_wilder_310\nenv = jo_wilder_310.make_env()\nm.predict(env)\n</code></pre>",
      "rawMarkdown": "I wanted to make simple baselines. The first baseline is using the average of correct answers for each question and finding a threshold, scoring .659. For my second baseline, I wanted to incorporate some more information. Considering the efficiency track of the competition, I thought \"what information can I get and use efficiently, ideally without even iterating through all the rows of data?\". My answer was that I can get the length of the test dataframe. I wanted to do a simple groupby... mean, but len(df) is not discrete enough. with chatgpt I found KBinsDiscretizer from sklearn and used that to bin len(df) feature so I could use groupby. \n\nHere is the code, scoring .671 public .675 private, currently 25th on efficiency leaderboard: \n```\nclass MeanKBD:\n    def __init__(self): \n        pass\n        \n    def fit(self, trn, trn_labels, raw=None):\n        if raw is not None: \n            raw = Path(raw)\n            trn = pd.read_csv(raw/'train.csv')\n            trn_labels = pd.read_csv(raw/'train_labels.csv')\n        trn_labels = add_basic_cols_to_ss(trn_labels)\n        tl = trn_labels.set_index(['session', 'session_level'])\n        trn['session_level'] = trn.level_group.map({'0-4': 0, '5-12': 1, '13-22': 2})\n        tl['row_count'] = trn.groupby(['session_id', 'session_level'])['session_level'].count()\n        \n        ##### Discretize length of session ########\n        kbd_dict = {}\n        for level in [0, 1, 2]: \n            tll = tl[tl.index.get_level_values(1) == level]\n            tll.row_count.values.reshape(-1, 1).shape\n            kbd = KBinsDiscretizer(n_bins=10)\n            tll = tll.assign(kbd=kbd.fit_transform(tll.row_count.values.reshape(-1, 1)).argmax(axis=1).astype(int))\n            tl.loc[tl.index.get_level_values(1) == level, 'kbd'] = tll['kbd']\n            kbd_dict[level] = kbd\n        tl['kbd'] = tl['kbd'].astype(int)\n        self.kbd_dict = kbd_dict\n        self.q_kbd_dict = tl.groupby(['q', 'kbd'])['correct'].mean().to_dict()\n        \n        ##### Find best threshold #####\n        prob = tl.apply(lambda row: (row['q'], row['kbd']), axis=1).map(self.q_kbd_dict)\n        best_score = -1\n        thresholds, scores = [], []\n        for th in np.arange(.2, .9, .01): \n            score = f1_score(tl['correct'], (prob > th).astype(int), average='macro')\n            thresholds.append(th)\n            scores.append(score)\n            if score > best_score: \n                best_score = score\n                best_th = th\n        print(f'best score: {best_score}, best_th {best_th}')\n        # sns.lineplot(x=thresholds, y=scores, marker='.')\n        # plt.show()\n        self.best_th = best_th\n        \n    def predict(self, env, fast=False): \n        if fast: \n            tl = env.test_labels.set_index(['session', 'session_level'])\n            trn = env.test\n            tl['row_count'] = trn.groupby(['session_id', 'session_level'])['session_level'].count()\n            for level in [0, 1, 2]:\n                mask = tl.index.get_level_values(1) == level\n                kbd = self.kbd_dict[level].transform(tl.loc[mask, 'row_count'].values.reshape(-1, 1))\n                tl.loc[mask, 'kbd'] = kbd.argmax(axis=1)\n            prob = tl.apply(lambda row: (row['q'], row['kbd']), axis=1).map(self.q_kbd_dict)\n            tl['correct'] = (prob > self.best_th).astype(int)\n            env.ss = tl\n            return None\n            \n        for i, (test, ss) in enumerate(env.iter_test()): \n            tmp = ss.copy()\n            tmp['row_count'] = len(test)\n            tmp['q'] = tmp.session_id.str.split('_').str[-1].str[1:].astype(int)\n            tmp['kbd'] = self.kbd_dict[i % 3].transform(tmp['row_count'].values.reshape(-1, 1)).argmax(axis=1)\n            prob = tmp.apply(lambda row: (row['q'], row['kbd']), axis=1).map(self.q_kbd_dict)\n            \n            ss['correct'] = (prob > self.best_th).astype(int)\n            env.predict(ss)\n\nimport joblib\nfrom gameplay import MeanKBD\n\nm = MeanKBD()\nm.fit(None, None, raw='/kaggle/input/gameplay-data')\njoblib.dump(m, 'm')\n\n###################### Put the rest of the code in an inference notebook ################\nimport joblib\nfrom gameplay import MeanKBD\n\nm = joblib.load('/kaggle/input/gameplay-mean-kbd-train/m')\nimport jo_wilder_310\nenv = jo_wilder_310.make_env()\nm.predict(env)\n```",
      "votes": null
    },
    {
      "id": "2326103",
      "postDate": "07/01/2023 22:48:15",
      "content": "<p>Thanks for sharing.</p>\n<p>Am I correct that your approach essentially uses one feature that is a number of events for a given session?</p>\n<p>Also, what was the inference time? I assume a couple of minutes?</p>",
      "rawMarkdown": "Thanks for sharing.\n\nAm I correct that your approach essentially uses one feature that is a number of events for a given session?\n\nAlso, what was the inference time? I assume a couple of minutes?",
      "votes": null
    },
    {
      "id": "2326390",
      "postDate": "07/02/2023 05:31:13",
      "content": "<p>Inference time just a few minutes. </p>\n<p>2 features: question number and test length. </p>\n<p>Algorithm: group by features and apply mean. </p>\n<p>For each set of questions, I count the number of events in the corresponding chapter of gameplay, which is given by the length of the test data frame given for that set of questions. This feature has many different values so I squish the ranges into just 10 values. Now I can group by question number and test length and take the mean. </p>",
      "rawMarkdown": "Inference time just a few minutes. \n\n2 features: question number and test length. \n\nAlgorithm: group by features and apply mean. \n\nFor each set of questions, I count the number of events in the corresponding chapter of gameplay, which is given by the length of the test data frame given for that set of questions. This feature has many different values so I squish the ranges into just 10 values. Now I can group by question number and test length and take the mean.",
      "votes": null
    },
    {
      "id": "2327598",
      "postDate": "07/03/2023 04:20:25",
      "content": "<p>Still hard for me to understand. But reading the code and what you wrote, and trying to restate it, is it like the below?</p>\n<p>So for a given question and filter/group by session_id and session_level, given all that, it's sort and bucket by row_count, splitting the train data into 10 buckets from smallest length bucket up to largest length bucket. For each bucket, take the percentage (aka mean) of the 'correct' label for that question. The when predicting, use row_count and question, determine the bucket, and map it to the percentage of that bucket.</p>\n<p>Finally, determine the best threshold, and replace the percentages with 0 or 1 based on above vs below that threshold.</p>",
      "rawMarkdown": "Still hard for me to understand. But reading the code and what you wrote, and trying to restate it, is it like the below?\n\nSo for a given question and filter/group by session_id and session_level, given all that, it's sort and bucket by row_count, splitting the train data into 10 buckets from smallest length bucket up to largest length bucket. For each bucket, take the percentage (aka mean) of the 'correct' label for that question. The when predicting, use row_count and question, determine the bucket, and map it to the percentage of that bucket.\n\nFinally, determine the best threshold, and replace the percentages with 0 or 1 based on above vs below that threshold.",
      "votes": null
    },
    {
      "id": "2328480",
      "postDate": "07/03/2023 15:58:00",
      "content": "<p>Thank you for sharing that snippet/baselines.</p>",
      "rawMarkdown": "Thank you for sharing that snippet/baselines.",
      "votes": null
    },
    {
      "id": "2328493",
      "postDate": "07/03/2023 16:16:50",
      "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> yes that is correct. The only small note is that I think the smallest and largest length bucket technically don't affect the bin cutoffs. The endpoint cutoffs will be where the 10% and 90% quantiles are. This way -inf and +inf will be in the first in last bucket during inference. There are other methods for choosing the bin cutoffs, but I think quantiles (evenly distributed) are the default for sklearns KBinsDiscretizer. </p>\n<p>Also for clarity, I trained 3 KBinsDiscretizers. One for each level_group.</p>",
      "rawMarkdown": "roberthatch yes that is correct. The only small note is that I think the smallest and largest length bucket technically don't affect the bin cutoffs. The endpoint cutoffs will be where the 10% and 90% quantiles are. This way -inf and +inf will be in the first in last bucket during inference. There are other methods for choosing the bin cutoffs, but I think quantiles (evenly distributed) are the default for sklearns KBinsDiscretizer. \n\nAlso for clarity, I trained 3 KBinsDiscretizers. One for each level_group.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2326103,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "07/01/2023 22:48:15",
      "content": "<p>Thanks for sharing.</p>\n<p>Am I correct that your approach essentially uses one feature that is a number of events for a given session?</p>\n<p>Also, what was the inference time? I assume a couple of minutes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2326390,
          "author_name": "chrisrichardmiles",
          "author_url": "",
          "post_date": "07/02/2023 05:31:13",
          "content": "<p>Inference time just a few minutes. </p>\n<p>2 features: question number and test length. </p>\n<p>Algorithm: group by features and apply mean. </p>\n<p>For each set of questions, I count the number of events in the corresponding chapter of gameplay, which is given by the length of the test data frame given for that set of questions. This feature has many different values so I squish the ranges into just 10 values. Now I can group by question number and test length and take the mean. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2327598,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "07/03/2023 04:20:25",
              "content": "<p>Still hard for me to understand. But reading the code and what you wrote, and trying to restate it, is it like the below?</p>\n<p>So for a given question and filter/group by session_id and session_level, given all that, it's sort and bucket by row_count, splitting the train data into 10 buckets from smallest length bucket up to largest length bucket. For each bucket, take the percentage (aka mean) of the 'correct' label for that question. The when predicting, use row_count and question, determine the bucket, and map it to the percentage of that bucket.</p>\n<p>Finally, determine the best threshold, and replace the percentages with 0 or 1 based on above vs below that threshold.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2328493,
                  "author_name": "chrisrichardmiles",
                  "author_url": "",
                  "post_date": "07/03/2023 16:16:50",
                  "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> yes that is correct. The only small note is that I think the smallest and largest length bucket technically don't affect the bin cutoffs. The endpoint cutoffs will be where the 10% and 90% quantiles are. This way -inf and +inf will be in the first in last bucket during inference. There are other methods for choosing the bin cutoffs, but I think quantiles (evenly distributed) are the default for sklearns KBinsDiscretizer. </p>\n<p>Also for clarity, I trained 3 KBinsDiscretizers. One for each level_group.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2328480,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "07/03/2023 15:58:00",
      "content": "<p>Thank you for sharing that snippet/baselines.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2325864": "I wanted to make simple baselines. The first baseline is using the average of correct answers for each question and finding a threshold, scoring .659. For my second baseline, I wanted to incorporate some more information. Considering the efficiency track of the competition, I thought \"what information can I get and use efficiently, ideally without even iterating through all the rows of data?\". My answer was that I can get the length of the test dataframe. I wanted to do a simple groupby... mean, but len(df) is not discrete enough. with chatgpt I found KBinsDiscretizer from sklearn and used that to bin len(df) feature so I could use groupby. \n\nHere is the code, scoring .671 public .675 private, currently 25th on efficiency leaderboard: \n```\nclass MeanKBD:\n    def __init__(self): \n        pass\n        \n    def fit(self, trn, trn_labels, raw=None):\n        if raw is not None: \n            raw = Path(raw)\n            trn = pd.read_csv(raw/'train.csv')\n            trn_labels = pd.read_csv(raw/'train_labels.csv')\n        trn_labels = add_basic_cols_to_ss(trn_labels)\n        tl = trn_labels.set_index(['session', 'session_level'])\n        trn['session_level'] = trn.level_group.map({'0-4': 0, '5-12': 1, '13-22': 2})\n        tl['row_count'] = trn.groupby(['session_id', 'session_level'])['session_level'].count()\n        \n        ##### Discretize length of session ########\n        kbd_dict = {}\n        for level in [0, 1, 2]: \n            tll = tl[tl.index.get_level_values(1) == level]\n            tll.row_count.values.reshape(-1, 1).shape\n            kbd = KBinsDiscretizer(n_bins=10)\n            tll = tll.assign(kbd=kbd.fit_transform(tll.row_count.values.reshape(-1, 1)).argmax(axis=1).astype(int))\n            tl.loc[tl.index.get_level_values(1) == level, 'kbd'] = tll['kbd']\n            kbd_dict[level] = kbd\n        tl['kbd'] = tl['kbd'].astype(int)\n        self.kbd_dict = kbd_dict\n        self.q_kbd_dict = tl.groupby(['q', 'kbd'])['correct'].mean().to_dict()\n        \n        ##### Find best threshold #####\n        prob = tl.apply(lambda row: (row['q'], row['kbd']), axis=1).map(self.q_kbd_dict)\n        best_score = -1\n        thresholds, scores = [], []\n        for th in np.arange(.2, .9, .01): \n            score = f1_score(tl['correct'], (prob > th).astype(int), average='macro')\n            thresholds.append(th)\n            scores.append(score)\n            if score > best_score: \n                best_score = score\n                best_th = th\n        print(f'best score: {best_score}, best_th {best_th}')\n        # sns.lineplot(x=thresholds, y=scores, marker='.')\n        # plt.show()\n        self.best_th = best_th\n        \n    def predict(self, env, fast=False): \n        if fast: \n            tl = env.test_labels.set_index(['session', 'session_level'])\n            trn = env.test\n            tl['row_count'] = trn.groupby(['session_id', 'session_level'])['session_level'].count()\n            for level in [0, 1, 2]:\n                mask = tl.index.get_level_values(1) == level\n                kbd = self.kbd_dict[level].transform(tl.loc[mask, 'row_count'].values.reshape(-1, 1))\n                tl.loc[mask, 'kbd'] = kbd.argmax(axis=1)\n            prob = tl.apply(lambda row: (row['q'], row['kbd']), axis=1).map(self.q_kbd_dict)\n            tl['correct'] = (prob > self.best_th).astype(int)\n            env.ss = tl\n            return None\n            \n        for i, (test, ss) in enumerate(env.iter_test()): \n            tmp = ss.copy()\n            tmp['row_count'] = len(test)\n            tmp['q'] = tmp.session_id.str.split('_').str[-1].str[1:].astype(int)\n            tmp['kbd'] = self.kbd_dict[i % 3].transform(tmp['row_count'].values.reshape(-1, 1)).argmax(axis=1)\n            prob = tmp.apply(lambda row: (row['q'], row['kbd']), axis=1).map(self.q_kbd_dict)\n            \n            ss['correct'] = (prob > self.best_th).astype(int)\n            env.predict(ss)\n\nimport joblib\nfrom gameplay import MeanKBD\n\nm = MeanKBD()\nm.fit(None, None, raw='/kaggle/input/gameplay-data')\njoblib.dump(m, 'm')\n\n###################### Put the rest of the code in an inference notebook ################\nimport joblib\nfrom gameplay import MeanKBD\n\nm = joblib.load('/kaggle/input/gameplay-mean-kbd-train/m')\nimport jo_wilder_310\nenv = jo_wilder_310.make_env()\nm.predict(env)\n```",
    "2326103": "Thanks for sharing.\n\nAm I correct that your approach essentially uses one feature that is a number of events for a given session?\n\nAlso, what was the inference time? I assume a couple of minutes?",
    "2326390": "Inference time just a few minutes. \n\n2 features: question number and test length. \n\nAlgorithm: group by features and apply mean. \n\nFor each set of questions, I count the number of events in the corresponding chapter of gameplay, which is given by the length of the test data frame given for that set of questions. This feature has many different values so I squish the ranges into just 10 values. Now I can group by question number and test length and take the mean.",
    "2327598": "Still hard for me to understand. But reading the code and what you wrote, and trying to restate it, is it like the below?\n\nSo for a given question and filter/group by session_id and session_level, given all that, it's sort and bucket by row_count, splitting the train data into 10 buckets from smallest length bucket up to largest length bucket. For each bucket, take the percentage (aka mean) of the 'correct' label for that question. The when predicting, use row_count and question, determine the bucket, and map it to the percentage of that bucket.\n\nFinally, determine the best threshold, and replace the percentages with 0 or 1 based on above vs below that threshold.",
    "2328480": "Thank you for sharing that snippet/baselines.",
    "2328493": "roberthatch yes that is correct. The only small note is that I think the smallest and largest length bucket technically don't affect the bin cutoffs. The endpoint cutoffs will be where the 10% and 90% quantiles are. This way -inf and +inf will be in the first in last bucket during inference. There are other methods for choosing the bin cutoffs, but I think quantiles (evenly distributed) are the default for sklearns KBinsDiscretizer. \n\nAlso for clarity, I trained 3 KBinsDiscretizers. One for each level_group."
  },
  "source": "meta"
}