{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Based on the Tensorflow Keras baseline notebook from https://www.kaggle.com/code/nadare/tensorflow-baseline-lb-0-681\n\n- Configurable hyperparameters are placed on top for easy tuning.\n- Seed is added for reproducible results\n\n--------------------\n\n### For CSC532/621 Machine Learning class:\n- Take note of the differences between this current notebook and the source\n- Document your notebook thoroughly (use ChatGPT/search engines as necessary but put the content in quotations and cite the source)\n- Only share and discuss with your teammates who are officially on your Kaggle team\n- Use the following to keep track of your runs (feel free to add additional columns):\n\n## Changelog\n\n|Version | Description | LB | Note |\n| --: | -- | -- | -- |\n| V0 | Soure [Notebook](https://www.kaggle.com/code/nadare/tensorflow-baseline-lb-0-681) | `0.681` |baseline (TF Keras DCN V2)|\n| V1 |Some finetuning of hyperparameters and added seed for reproducible results | `0.682` ||\n| V2 |Added 1 num_linear layer and used the original 'gelu' activation function  | `0.682` |LB slightly lower than V1|\n| V3 |Quick save |  ||\n| V4 |Quick save |  |changed title to reflect CSC532 Machine Learning class|\n","metadata":{}},{"cell_type":"code","source":"import gc\nimport os\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nimport tensorflow_addons as tfa\nimport random\n\nfrom tqdm.notebook import tqdm","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:18:49.457013Z","start_time":"2023-02-19T14:18:47.989756Z"},"execution":{"iopub.status.busy":"2023-02-20T10:51:28.799473Z","iopub.execute_input":"2023-02-20T10:51:28.799933Z","iopub.status.idle":"2023-02-20T10:51:37.162875Z","shell.execute_reply.started":"2023-02-20T10:51:28.799842Z","shell.execute_reply":"2023-02-20T10:51:37.161667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code imports necessary libraries such as gc, os, numpy, pandas, tensorflow, tensorflow_addons, and random. It also imports the tqdm library from the notebook submodule for displaying progress bars during loops.\n\nThe gc library is used for garbage collection and memory management. The os library provides a way to interact with the operating system, such as getting the current working directory. numpy and pandas are popular libraries for working with arrays and dataframes, respectively. tensorflow is a popular deep learning library, and tensorflow_addons provides additional functionality for TensorFlow. random is a library for generating random numbers.","metadata":{}},{"cell_type":"code","source":"# Config\nEPOCHS = 20 # baseline was 10\nSEED = 42 # baseline model was stochastic\nthreshold = 100\nLR = 1.e-3\nepsilon = 1.e-9\n\nFEAT_DIM = 8\nOUT_DIM = 32\nNUM_CROSS = 5\nNUM_LINEAR = 1 # baseline was 0\nBATCH = 128\nACT_FUN = \"gelu\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines several constant variables used for the model configuration:\n\nEPOCHS: The number of epochs to train the model.\nSEED: The random seed to use for reproducibility.\nthreshold: A threshold value for the r2_score metric to determine whether to save the model weights or not.\nLR: The learning rate for the Adam optimizer.\nepsilon: A small value to prevent division by zero in the Adam optimizer.\nFEAT_DIM: The number of dimensions for the embeddings of categorical features.\nOUT_DIM: The output dimension of the MLP head.\nNUM_CROSS: The number of layers for the cross network.\nNUM_LINEAR: The number of fully connected layers before the MLP head.\nBATCH: The batch size for training.\nACT_FUN: The activation function to use for the MLP layers. In this case, it is set to \"gelu\".","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# Seed all random number generators\ndef seed_everything(seed=SEED):\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    random.seed(seed)\n    np.random.seed(seed)\n    tf.random.set_seed(seed)\n\nseed_everything()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines a function seed_everything() which sets the random seeds for the Python, NumPy, and TensorFlow libraries. This is done to ensure that the results are reproducible.\n\nThe function takes an argument seed, which is set to the value of SEED by default. The os.environ['PYTHONHASHSEED'] line sets the Python hash seed, which affects the order in which items are inserted into dictionaries, to the value of seed. The random.seed(seed) line sets the seed for the random module. The np.random.seed(seed) line sets the seed for NumPy's random number generator. Finally, the tf.random.set_seed(seed) line sets the seed for TensorFlow's random number generator. All of these steps together ensure that the random number generators used by the model are initialized with the same seed each time the code is run, leading to consistent results.","metadata":{}},{"cell_type":"code","source":"feature_columns = [\"event_name\", \"name\", \"level\", \"fqid\", \"room_fqid\", \"text_fqid\", \"text\"]\nfeature_ix_columns = [col + \"_ix\" for col in feature_columns]\nusecols = [\"session_id\", \"level_group\"] + feature_columns\n\ntrain_df = pd.read_csv(\"../input/predict-student-performance-from-game-play/train.csv\", usecols=usecols)","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:19:02.846725Z","start_time":"2023-02-19T14:18:49.458018Z"},"execution":{"iopub.status.busy":"2023-02-20T10:51:37.164438Z","iopub.execute_input":"2023-02-20T10:51:37.165012Z","iopub.status.idle":"2023-02-20T10:52:39.313284Z","shell.execute_reply.started":"2023-02-20T10:51:37.16498Z","shell.execute_reply":"2023-02-20T10:52:39.311941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code reads in a CSV file called \"train.csv\" using Pandas' read_csv() function. The file is assumed to be located in the \"../input/predict-student-performance-from-game-play/\" directory. The usecols parameter specifies which columns to read from the file - in this case, it reads the \"session_id\", \"level_group\", \"event_name\", \"name\", \"level\", \"fqid\", \"room_fqid\", \"text_fqid\", and \"text\" columns.\n\nThe data is read into a Pandas DataFrame called train_df. The feature columns are the categorical features used for the model, and feature_ix_columns is a list of the feature columns with the \"_ix\" suffix added to indicate that they will be used for embedding. The train_df DataFrame will be used for training the model.","metadata":{}},{"cell_type":"code","source":"# category to int\nfeature_dicts = {}\nfor col in tqdm(feature_columns):\n    vc = train_df[col].fillna(\"nan_value\").astype(str).value_counts()\n    vc = vc[vc.values >= threshold]\n    feature_dict = {k: i for i, k in enumerate(vc.index, start=1)}\n    train_df[col + \"_ix\"] = np.vectorize(lambda x: feature_dict.get(x, 0))(train_df[col].fillna(\"nan_value\").astype(str)).astype(np.int32)\n    feature_dicts[col] = feature_dict\n    del train_df[col]\n    gc.collect()\nfeature_num_vocabs = [len(feature_dicts[k]) + 1 for k in feature_columns]\n\nfor col, num_vocab in zip(feature_columns, feature_num_vocabs):\n    print(col, num_vocab)","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:19:24.566707Z","start_time":"2023-02-19T14:19:02.847734Z"},"execution":{"iopub.status.busy":"2023-02-20T10:52:39.314507Z","iopub.execute_input":"2023-02-20T10:52:39.314898Z","iopub.status.idle":"2023-02-20T10:53:42.658977Z","shell.execute_reply.started":"2023-02-20T10:52:39.314869Z","shell.execute_reply":"2023-02-20T10:53:42.653998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code creates integer mappings for the categorical features in the train_df DataFrame. For each feature column in feature_columns, it counts the number of occurrences of each unique value in the column, fills missing values with \"nan_value\", and retains only the values with counts greater than or equal to threshold. It then creates a dictionary that maps each value to an integer index, starting from 1. The 0 index is reserved for out-of-vocabulary values.\n\nThe code then applies the integer mapping to each value in the feature column by calling np.vectorize() on a lambda function that uses the corresponding feature dictionary to map each value to its integer index. The resulting integer-mapped column is added to train_df with the \"_ix\" suffix, and the original feature column is deleted from train_df to save memory. The gc.collect() line explicitly triggers garbage collection to free up memory.\n\nFinally, the code prints the number of distinct values for each feature column, including the out-of-vocabulary index. These values will be used to determine the number of categories in each embedding layer of the model.","metadata":{}},{"cell_type":"code","source":"train_df[\"level_group\"] = train_df[\"level_group\"].map({k: v for v, k in enumerate(['0-4', '5-12', '13-22'])})","metadata":{"execution":{"iopub.status.busy":"2023-02-20T10:53:42.662465Z","iopub.execute_input":"2023-02-20T10:53:42.663174Z","iopub.status.idle":"2023-02-20T10:53:43.986349Z","shell.execute_reply.started":"2023-02-20T10:53:42.663127Z","shell.execute_reply":"2023-02-20T10:53:43.985106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code maps the string values in the \"level_group\" column of the train_df DataFrame to integers. The original values in \"level_group\" are assumed to be one of the strings '0-4', '5-12', or '13-22'. The map() method is used to apply a dictionary that maps these strings to integers. Specifically, the dictionary used is {k: v for v, k in enumerate(['0-4', '5-12', '13-22'])}, which creates a dictionary that maps each string in the list to an integer index, starting from 0. The resulting integer-mapped column is saved back to the \"level_group\" column in train_df. This mapping is done to convert the categorical \"level_group\" feature into a numerical feature that can be used by the model.","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T10:53:43.988087Z","iopub.execute_input":"2023-02-20T10:53:43.988436Z","iopub.status.idle":"2023-02-20T10:53:44.00339Z","shell.execute_reply.started":"2023-02-20T10:53:43.988405Z","shell.execute_reply":"2023-02-20T10:53:44.002395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code outputs information about the train_df DataFrame using the info() method. This method prints a summary of the DataFrame's data types, column names, number of non-null values, and memory usage.\n\nThe output will include the number of rows in the DataFrame, the names of each column, the number of non-null values in each column, and the data type of each column. This information can be useful for identifying missing values or unexpected data types that need to be handled before training the model.","metadata":{}},{"cell_type":"code","source":"# aggregate unique_feature\nunique_train_df = train_df[feature_ix_columns].drop_duplicates()\nunique_train_df[\"unique_ix\"] = np.arange(unique_train_df.shape[0])\n\ntrain_df = train_df.merge(unique_train_df,\n                          on=feature_ix_columns,\n                          how=\"left\")\nunique_feature = tf.convert_to_tensor(unique_train_df[feature_ix_columns].values)\n\nfor col in feature_ix_columns:\n    del train_df[col]\n    gc.collect()\n\nunique_feature.shape","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:19:29.917509Z","start_time":"2023-02-19T14:19:24.568708Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:44.004716Z","iopub.execute_input":"2023-02-20T10:53:44.005319Z","iopub.status.idle":"2023-02-20T10:53:51.021492Z","shell.execute_reply.started":"2023-02-20T10:53:44.005282Z","shell.execute_reply":"2023-02-20T10:53:51.020661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code aggregates unique values of the integer-mapped categorical features in the train_df DataFrame. It first selects the columns in train_df that correspond to the integer-mapped feature columns using train_df[feature_ix_columns]. It then drops any duplicate rows in the resulting DataFrame using drop_duplicates().\n\nThe code creates a new column in the resulting DataFrame that assigns a unique integer index to each unique combination of feature values using np.arange(unique_train_df.shape[0]). The resulting integer index is added to unique_train_df as a new column called \"unique_ix\".\n\nThe code then merges the \"unique_ix\" column from unique_train_df back into train_df using merge(). The merge is done on the integer-mapped feature columns specified by feature_ix_columns. The resulting DataFrame has a new column called \"unique_ix\" that contains the unique integer index for each row.\n\nThe code creates a TensorFlow tensor called \"unique_feature\" that contains the integer-mapped feature values for each unique combination of features. This tensor will be used to initialize the embedding layer weights.\n\nFinally, the code deletes the integer-mapped feature columns from train_df and runs garbage collection to free up memory.\n\nThe output of the last line shows the shape of the unique_feature tensor, which should be (num_unique_features, num_categorical_features), where num_unique_features is the number of unique combinations of features and num_categorical_features is the number of categorical features.","metadata":{}},{"cell_type":"code","source":"# session to index\ntrain_session_map = {k: i for i, k in enumerate(train_df[\"session_id\"].unique())}\ntrain_df[\"session_ix\"] = train_df[\"session_id\"].map(train_session_map)\nnum_session = train_df[\"session_ix\"].max() + 1\nnum_session","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:19:30.105616Z","start_time":"2023-02-19T14:19:29.91951Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:51.022892Z","iopub.execute_input":"2023-02-20T10:53:51.023411Z","iopub.status.idle":"2023-02-20T10:53:51.349414Z","shell.execute_reply.started":"2023-02-20T10:53:51.023378Z","shell.execute_reply":"2023-02-20T10:53:51.348346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code maps each unique session ID in the train_df DataFrame to a unique integer index using a Python dictionary comprehension. The resulting dictionary is stored in the train_session_map variable.\n\nThe code then creates a new column in train_df called \"session_ix\" that maps each session ID to its corresponding integer index using the map() method and the train_session_map dictionary. The integer index for each session is determined by its position in the unique session ID list.\n\nThe code calculates the total number of sessions by finding the maximum integer index in the \"session_ix\" column and adding 1. This number is stored in the num_session variable.\n\nThe resulting integer index for each session will be used to group sessions together during training and validation.","metadata":{}},{"cell_type":"code","source":"# history to ragged tensor\ntrain_histories = []\nfor i in range(3):\n    tmp_df = train_df[[\"session_ix\", \"unique_ix\", \"level_group\"]].query(f\"level_group == {i}\")\n    train_histories.append(tf.RaggedTensor.from_value_rowids(values=tf.convert_to_tensor(tmp_df[\"unique_ix\"].astype(np.int64)),\n                                                             value_rowids=tf.convert_to_tensor(tmp_df[\"session_ix\"].astype(np.int64)),\n                                                             nrows=num_session))","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:19:31.689464Z","start_time":"2023-02-19T14:19:30.106617Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:51.351121Z","iopub.execute_input":"2023-02-20T10:53:51.351795Z","iopub.status.idle":"2023-02-20T10:53:53.288731Z","shell.execute_reply.started":"2023-02-20T10:53:51.351735Z","shell.execute_reply":"2023-02-20T10:53:53.287722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code creates a list of three RaggedTensor objects, train_histories, one for each of the three possible values of \"level_group\" (0, 1, or 2). Each RaggedTensor represents the training history for sessions in the corresponding level group.\n\nThe RaggedTensor is created using the tf.RaggedTensor.from_value_rowids() method, which constructs a ragged tensor from a 1D tensor of values and a 1D tensor of row indices. In this case, the values tensor contains the unique feature indices for each session, and the row indices tensor contains the integer indices of each session.\n\nThe values and value_rowids arguments are created using tf.convert_to_tensor() to convert the corresponding columns of a filtered subset of the train_df DataFrame to tensors of int64 data type. The nrows argument is set to the total number of sessions, num_session, which ensures that all RaggedTensor objects have the same number of rows.","metadata":{}},{"cell_type":"code","source":"del train_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T10:53:53.29196Z","iopub.execute_input":"2023-02-20T10:53:53.292324Z","iopub.status.idle":"2023-02-20T10:53:53.487997Z","shell.execute_reply.started":"2023-02-20T10:53:53.292292Z","shell.execute_reply":"2023-02-20T10:53:53.487016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These two lines of code delete the train_df DataFrame and run garbage collection to free up memory. Since the train_df DataFrame is no longer needed, it can be deleted to free up memory for subsequent operations. The gc.collect() function is called to ensure that any unreferenced objects in memory are promptly garbage-collected to free up memory.","metadata":{}},{"cell_type":"code","source":"# label to tensor\ntrain_label_df = pd.read_csv(\"../input/predict-student-performance-from-game-play/train_labels.csv\")\ntrain_label_df[\"session_ix\"] = train_label_df[\"session_id\"].map(lambda x: train_session_map[int(x.split(\"_\")[0])])\ntrain_label_df[\"question_ix\"] = train_label_df[\"session_id\"].map(lambda x: int(x.split(\"_\")[1][1:]))\n\ntrain_label = train_label_df.pivot_table(columns=\"question_ix\",\n                                         index=\"session_ix\",\n                                         values=\"correct\",\n                                         aggfunc=\"first\").loc[np.arange(num_session)].values.astype(np.float32)\ntrain_labels = [tf.convert_to_tensor(train_label[:, :3]), tf.convert_to_tensor(train_label[:, 3:13]), tf.convert_to_tensor(train_label[:, 13:])]\n\ndel train_label_df\ngc.collect()","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:19:31.953894Z","start_time":"2023-02-19T14:19:31.690973Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:53.491576Z","iopub.execute_input":"2023-02-20T10:53:53.493826Z","iopub.status.idle":"2023-02-20T10:53:54.446288Z","shell.execute_reply.started":"2023-02-20T10:53:53.493773Z","shell.execute_reply":"2023-02-20T10:53:54.445486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These lines of code read the training labels from the \"../input/predict-student-performance-from-game-play/train_labels.csv\" file, map the session IDs to session indices, map the question IDs to question indices, and create a pivot table of correct answers. The train_labels variable is then created as a list of three tensors, with each tensor containing the correct answers for a particular subset of questions. Finally, the train_label_df DataFrame is deleted and garbage collection is run to free up memory.","metadata":{}},{"cell_type":"code","source":"# define dataset\nclass DataLoader():\n\n    def __init__(self, train_histories, train_labels, train_unique_feataures):\n        self.train_histories = train_histories\n        self.train_labels = train_labels\n        self.train_unique_feataures = train_unique_feataures\n\n    def call(self, inputs):\n        session_ix = inputs\n        raw_histories = [tf.gather(self.train_histories[i], session_ix) for i in range(3)]\n        historiy_lengths = [tf.shape(h.values)[0] for h in raw_histories]\n        unique_ix, unique_idx = tf.unique(tf.concat([h.values for h in raw_histories], axis=0))\n        history_values = tf.split(unique_idx, historiy_lengths, axis=0)\n        histories = [tf.RaggedTensor.from_value_rowids(values=history_values[i],\n                                                       value_rowids=raw_histories[i].value_rowids(),\n                                                       nrows=raw_histories[i].nrows()) for i in range(3)]\n        inputs = {}\n        for i in range(3):\n            inputs[f\"history_{i}\"] = histories[i]\n            label = tf.gather(self.train_labels[i], session_ix)\n            inputs[f\"label_{i}\"] = label\n        inputs[\"unique_feature\"] = tf.gather(self.train_unique_feataures, unique_ix)\n        return inputs\n\n\n# define model\nclass DCNV2Model(tf.keras.Model):\n\n    def __init__(self, feature_num_vocabs, feat_dim, out_dim, num_cross, num_linear):\n        super(DCNV2Model, self).__init__()\n        self.num_features = len(feature_num_vocabs)\n\n        self.feature_num_vocabs = feature_num_vocabs\n        self.feat_dim = feat_dim\n        self.out_dim = out_dim\n        self.num_cross = num_cross\n        self.num_linear = num_linear\n\n        self.input_dim = feat_dim * self.num_features\n        self.embedding_layers = [tf.keras.layers.Embedding(feature_num_vocabs[i], feat_dim) for i in range(self.num_features)]\n        self.cross_in_layers = [tf.keras.layers.Dense(self.feat_dim) for _ in range(self.num_cross)]\n        self.cross_out_layers = [tf.keras.layers.Dense(self.input_dim) for _ in range(self.num_cross)]\n        self.linear_layers = [tf.keras.layers.Dense(self.input_dim, activation=ACT_FUN) for _ in range(self.num_linear)]\n        self.out_layer = tf.keras.layers.Dense(self.out_dim)\n        \n    def call(self, inputs):\n        X = []\n        for i in range(self.num_features):\n            X.append(self.embedding_layers[i](tf.gather(inputs, i, axis=1)))\n        X = tf.concat(X, axis=1)\n        X0 = tf.identity(X)\n\n        for i in range(self.num_cross):\n            X = X0 * self.cross_out_layers[i](self.cross_in_layers[i](X)) + X\n\n        for i in range(self.num_linear):\n            X = self.linear_layers[i](X)\n\n        X = self.out_layer(X)\n        return X\n\nclass Predictor(tf.keras.Model):\n\n    def __init__(self, out_dim):\n        super(Predictor, self).__init__()\n        self.out_layer = tf.keras.layers.Dense(out_dim)\n\n    def call(self, inputs):\n        return self.out_layer(inputs)\n\n\nclass Trainer(tf.keras.Model):\n\n    def __init__(self, emb_model, predictors):\n        super(Trainer, self).__init__()\n        self.emb_model = emb_model\n        self.predictors = predictors\n        self.eps = epsilon\n\n    def call(self, inputs):\n        unique_emb = self.emb_model(inputs[\"unique_feature\"])\n        loss_sum = 0.\n        pos_true_positive = 0.\n        pos_false_positive = 0.\n        pos_false_negative = 0.\n        neg_true_positive = 0.\n        neg_false_positive = 0.\n        neg_false_negative = 0.\n        correct = 0.\n        for i in range(3):\n            pred_emb = tf.reduce_sum(tf.gather(unique_emb, inputs[f\"history_{i}\"]), axis=1)\n            pred_val = tf.clip_by_value(tf.math.sigmoid(self.predictors[i](tf.nn.l2_normalize(pred_emb, axis=1))), self.eps, 1. - self.eps)\n            # Binary Cross Entropy\n            loss = -inputs[f\"label_{i}\"] * tf.math.log(pred_val) - (1. - inputs[f\"label_{i}\"]) * tf.math.log(1. - pred_val)\n            loss_sum += tf.reduce_sum(tf.reduce_mean(loss, axis=0))\n\n            # F1-macro\n            pred_label = pred_val > .5\n            bool_label = inputs[f\"label_{i}\"] == 1.\n            correct += tf.cast(tf.math.count_nonzero(pred_label == bool_label), \"float32\")\n            pos_true_positive += tf.cast(tf.math.count_nonzero(tf.math.logical_and(bool_label, pred_label)), \"float32\")\n            pos_false_positive += tf.cast(tf.math.count_nonzero(tf.math.logical_and(tf.math.logical_not(bool_label), pred_label)), \"float32\")\n            pos_false_negative += tf.cast(tf.math.count_nonzero(tf.math.logical_and(bool_label, tf.math.logical_not(pred_label))), \"float32\")\n\n            \n            pred_label = pred_val < .5\n            bool_label = inputs[f\"label_{i}\"] == 0.\n            neg_true_positive += tf.cast(tf.math.count_nonzero(tf.math.logical_and(bool_label, pred_label)), \"float32\")\n            neg_false_positive += tf.cast(tf.math.count_nonzero(tf.math.logical_and(tf.math.logical_not(bool_label), pred_label)), \"float32\")\n            neg_false_negative += tf.cast(tf.math.count_nonzero(tf.math.logical_and(bool_label, tf.math.logical_not(pred_label))), \"float32\")\n\n            \n        accuracy = correct / tf.cast(tf.shape(inputs[f\"label_0\"])[0] * 18, \"float32\")\n        pos_recall = pos_true_positive / tf.maximum(self.eps, pos_true_positive + pos_false_negative)\n        pos_precision = pos_true_positive / tf.maximum(self.eps, pos_true_positive + pos_false_positive)\n        pos_f1 = 2*pos_recall*pos_precision / tf.maximum(self.eps, pos_recall + pos_precision)\n        \n        neg_recall = neg_true_positive / tf.maximum(self.eps, neg_true_positive + neg_false_negative)\n        neg_precision = neg_true_positive / tf.maximum(self.eps, neg_true_positive + neg_false_positive)\n        neg_f1 = 2*neg_recall*neg_precision / tf.maximum(self.eps, neg_recall + neg_precision)\n        \n        return loss_sum, (pos_f1 + neg_f1)/2., accuracy\n\n    def predict_proba(self, inputs):\n        unique_emb = self.emb_model(inputs[\"unique_feature\"])\n        labels = []\n        pred_vals = []\n        for i in range(3):\n            pred_emb = tf.reduce_sum(tf.gather(unique_emb, inputs[f\"history_{i}\"]), axis=1)\n            pred_val = tf.clip_by_value(tf.math.sigmoid(self.predictors[i](tf.nn.l2_normalize(pred_emb, axis=1))), self.eps, 1. - self.eps)\n            pred_vals.append(pred_val)\n            labels.append(inputs[f\"label_{i}\"])\n        return tf.concat(pred_vals, axis=1), tf.concat(labels, axis=1)\n","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:25:52.879446Z","start_time":"2023-02-19T14:25:52.86024Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:54.447878Z","iopub.execute_input":"2023-02-20T10:53:54.448178Z","iopub.status.idle":"2023-02-20T10:53:54.482521Z","shell.execute_reply.started":"2023-02-20T10:53:54.448151Z","shell.execute_reply":"2023-02-20T10:53:54.481455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code you provided is a TensorFlow implementation of a Deep & Cross Network (DCN) model, along with additional components for training and evaluation. The model is a type of neural network that combines a deep neural network (DNN) and a cross network to capture both low-order and high-order feature interactions.\n\nHere is a brief summary of the main components of the code:\n\nDataLoader: A class that loads and processes training data for the model. The input to the model consists of a set of historical feature values, along with binary labels indicating whether the user clicked on a certain item or not. The DataLoader converts these inputs into a format that can be processed by the model, including splitting the historical feature values into separate ragged tensors and embedding categorical features.\nDCNV2Model: The main model class, which implements the DCN architecture. The input features are passed through a series of embedding layers, cross network layers, and fully connected (linear) layers, before being fed into an output layer that produces a binary prediction for each input instance.\nPredictor: A separate model class that produces binary predictions for each of the three types of input features in the dataset. This is used during training to compute the loss and evaluate the model performance.\nTrainer: A class that implements the training loop for the model. Given a set of inputs and binary labels, it uses the DCNV2Model and Predictor classes to compute the loss and update the model weights.\nOverall, this code represents a fairly standard implementation of a DCN model for click-through rate prediction in recommender systems. The key innovation of this model is the use of cross network layers to capture high-order feature interactions, which can be more difficult to capture using a purely deep neural network.","metadata":{}},{"cell_type":"code","source":"# evaluate model\nfrom sklearn.model_selection import train_test_split\n\n# define model and dataset\ndata_loader = DataLoader(train_histories, train_labels, unique_feature)\nemb_model = DCNV2Model(feature_num_vocabs, feat_dim=FEAT_DIM, out_dim=OUT_DIM, num_cross=NUM_CROSS, num_linear=NUM_LINEAR)\npredictors = [Predictor(3), Predictor(10), Predictor(5)]\ntrainer = Trainer(emb_model, predictors)\n\noptimizer = tfa.optimizers.LazyAdam(learning_rate=LR)\nloss_metric = tf.keras.metrics.Mean()\nf1_metric = tf.keras.metrics.Mean()\nacc_metric = tf.keras.metrics.Mean()\n\ntrain_sessions = np.arange(num_session)\ndev_sessions, val_sessions = train_test_split(train_sessions, test_size=2000, shuffle=True)\ndev_dataset = tf.data.Dataset.from_tensor_slices(tf.convert_to_tensor(dev_sessions))\\\n                             .shuffle(num_session, reshuffle_each_iteration=True)\\\n                             .batch(BATCH)\\\n                             .map(data_loader.call)\\\n                             .prefetch(tf.data.AUTOTUNE)","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:25:53.553354Z","start_time":"2023-02-19T14:25:53.527075Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:54.484079Z","iopub.execute_input":"2023-02-20T10:53:54.484397Z","iopub.status.idle":"2023-02-20T10:53:55.589985Z","shell.execute_reply.started":"2023-02-20T10:53:54.484367Z","shell.execute_reply":"2023-02-20T10:53:55.588808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines the necessary components for evaluating a model.\n\nFirst, it imports the train_test_split function from the sklearn.model_selection module to split the data into training and validation sets.\n\nNext, it defines the model and dataset using the DataLoader, DCNV2Model, Predictor, and Trainer classes. The DataLoader is responsible for loading and processing the data, while the DCNV2Model is the actual model being trained. The Predictor is used to generate embeddings for each session, which are then used as inputs to the model. The Trainer is used to train the model.\n\nThe optimizer used is tfa.optimizers.LazyAdam, which is a variant of the Adam optimizer that uses lazy updates to improve performance.\n\nThe metrics used to evaluate the model are the mean loss, mean F1 score, and mean accuracy, which are all defined using tf.keras.metrics.Mean.\n\nFinally, the training set is split into development and validation sets using train_test_split, and the dev_dataset is created using tf.data.Dataset.from_tensor_slices to convert the development set into a TensorFlow dataset. The dataset is then shuffled and batched using shuffle, batch, and prefetch, respectively. The map function is used to apply the data_loader.call function to each batch of the dataset, which processes the data and generates the inputs for the model.","metadata":{}},{"cell_type":"code","source":"# training\n\n@tf.function(experimental_relax_shapes=True)\ndef forward_step(batch_inputs):\n    with tf.GradientTape() as tape:\n        loss, f1, acc = trainer(batch_inputs, training=True)\n    gradients = tape.gradient(loss, trainer.trainable_variables)\n    optimizer.apply_gradients(zip(gradients, trainer.trainable_variables))\n    return loss, f1, acc\n\nwith tf.device(\"CPU: 0\"):\n    for epoch in range(EPOCHS):\n        with tqdm(total=len(dev_dataset)) as pbar:\n            for batch_inputs in dev_dataset:\n                loss, f1, acc = forward_step(batch_inputs)\n                loss_metric(loss)\n                f1_metric(f1)\n                acc_metric(acc)\n                progress = {\"BCE\": loss_metric.result().numpy(), \"f1\": f1_metric.result().numpy(), \"accuracy\": acc_metric.result().numpy()}\n                pbar.set_postfix(progress)\n                pbar.update(1)\n        print(\"epoch:\", epoch+1)\n        print(\"dev:\", loss_metric.result().numpy(), f1_metric.result().numpy())\n        val_loss, val_f1, val_acc = trainer(data_loader.call(val_sessions))\n        print(\"val:\", val_loss.numpy(), val_f1.numpy(), val_acc.numpy())\n        loss_metric.reset_states()\n        f1_metric.reset_states()","metadata":{"ExecuteTime":{"end_time":"2023-02-19T14:27:10.314161Z","start_time":"2023-02-19T14:25:57.61061Z"},"execution":{"iopub.status.busy":"2023-02-20T10:53:55.592418Z","iopub.execute_input":"2023-02-20T10:53:55.592761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This appears to be the main training loop of the model. Here's what's happening:\n\n@tf.function decorator is used to make the function forward_step a TensorFlow computation graph, which can be optimized for execution on GPU or TPU.\nwith tf.GradientTape() as tape: starts a gradient tape, which is a TensorFlow object that records operations for automatic differentiation.\nloss, f1, acc = trainer(batch_inputs, training=True) computes the loss, f1 score, and accuracy of the model on the current batch of input data.\ngradients = tape.gradient(loss, trainer.trainable_variables) computes the gradients of the loss with respect to the trainable variables of the model.\noptimizer.apply_gradients(zip(gradients, trainer.trainable_variables)) applies the gradients to update the parameters of the model.\nwith tqdm(total=len(dev_dataset)) as pbar: creates a progress bar for the training loop.\nfor batch_inputs in dev_dataset: iterates through the batches of input data in the dev dataset.\nloss, f1, acc = forward_step(batch_inputs) performs a forward step of the model and computes the loss, f1 score, and accuracy.\nloss_metric(loss), f1_metric(f1), and acc_metric(acc) update the respective metrics with the values from the current batch.\npbar.set_postfix(progress) updates the progress bar with the current values of the metrics.\npbar.update(1) increments the progress bar by 1.\nAfter each epoch, the model is evaluated on the validation set, and the metrics are printed to the console. Then, the metrics are reset for the next epoch using loss_metric.reset_states() and f1_metric.reset_states().","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfrom sklearn.metrics import f1_score\ndev_pred, dev_label = trainer.predict_proba(data_loader.call(dev_sessions))\nval_pred, val_label = trainer.predict_proba(data_loader.call(val_sessions))\n\ndev_f1_scores = []\nval_f1_scores = []\nfor th in np.linspace(0.1, 0.9, 80):\n    dev_f1_scores.append(f1_score(dev_label.numpy().reshape(-1), (dev_pred > th).numpy().reshape(-1), average=\"macro\"))\n    val_f1_scores.append(f1_score(val_label.numpy().reshape(-1), (val_pred > th).numpy().reshape(-1), average=\"macro\"))\n\nplt.plot(np.linspace(0.1, 0.9, 80), dev_f1_scores, label=\"dev\")\nplt.plot(np.linspace(0.1, 0.9, 80), val_f1_scores, label=\"val\")\nplt.ylim(0, 1)\n\nprint(max(dev_f1_scores), max(val_f1_scores))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that you are trying to plot the F1 score for different thresholds on both the dev and validation sets. This can be a useful visualization to determine the optimal threshold for the model's binary classification output. The F1 score is the harmonic mean of precision and recall, and it is often used in imbalanced classification problems.\n\nThe code you provided looks good. You are calculating the F1 score for different thresholds using the f1_score function from scikit-learn, and then plotting the results using Matplotlib. You are also printing the maximum F1 scores obtained for both the dev and validation sets.","metadata":{}},{"cell_type":"code","source":"print(max(dev_f1_scores), max(val_f1_scores))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The maximum F1 score on the dev set is 0.43919920623477816, and the maximum F1 score on the validation set is 0.41007034415881984.","metadata":{}},{"cell_type":"code","source":"# use all_data for train\n# define model and dataset\ndata_loader = DataLoader(train_histories, train_labels, unique_feature)\nemb_model = DCNV2Model(feature_num_vocabs, feat_dim=FEAT_DIM, out_dim=OUT_DIM, num_cross=NUM_CROSS, num_linear=NUM_LINEAR)\npredictors = [Predictor(3), Predictor(10), Predictor(5)]\ntrainer = Trainer(emb_model, predictors)\n\noptimizer = tfa.optimizers.LazyAdam(learning_rate=LR)\nloss_metric = tf.keras.metrics.Mean()\nf1_metric = tf.keras.metrics.Mean()\nacc_metric = tf.keras.metrics.Mean()\n\ntrain_sessions = np.arange(num_session)\ntrain_dataset = tf.data.Dataset.from_tensor_slices(tf.convert_to_tensor(train_sessions))\\\n                               .shuffle(num_session, reshuffle_each_iteration=True)\\\n                               .batch(BATCH)\\\n                               .map(data_loader.call)\\\n                               .prefetch(tf.data.AUTOTUNE)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great! Now we can train the model on the entire dataset. You can use the same training loop as before, just replace dev_dataset with train_dataset and change the printing statements to show the training results instead of validation results.","metadata":{}},{"cell_type":"code","source":"# training\n\n@tf.function(experimental_relax_shapes=True)\ndef forward_step(batch_inputs):\n    with tf.GradientTape() as tape:\n        loss, f1, acc = trainer(batch_inputs, training=True)\n    gradients = tape.gradient(loss, trainer.trainable_variables)\n    optimizer.apply_gradients(zip(gradients, trainer.trainable_variables))\n    return loss, f1, acc\n\nwith tf.device(\"CPU: 0\"):\n    for epoch in range(EPOCHS):\n        with tqdm(total=len(train_dataset)) as pbar:\n            for batch_inputs in train_dataset:\n                loss, f1, acc = forward_step(batch_inputs)\n                loss_metric(loss)\n                f1_metric(f1)\n                acc_metric(acc)\n                progress = {\"BCE\": loss_metric.result().numpy(), \"f1\": f1_metric.result().numpy(), \"accuracy\": acc_metric.result().numpy()}\n                pbar.set_postfix(progress)\n                pbar.update(1)\n        print(\"epoch:\", epoch+1)\n        print(\"train:\", loss_metric.result().numpy(), f1_metric.result().numpy())\n        loss_metric.reset_states()\n        f1_metric.reset_states()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems like the model is being trained on the CPU instead of the GPU. Training on a GPU would usually result in a much faster training time.\n\nTo use the GPU, you can set the tf.device context to the GPU device index, like tf.device(\"/GPU:0\").\n\nAdditionally, the progress bar should be updated after each batch, not after each epoch. This would provide a better indication of the training progress.","metadata":{}},{"cell_type":"code","source":"train_pred, train_label = trainer.predict_proba(data_loader.call(train_sessions))\n\ntrain_f1_scores = []\nfor th in np.linspace(0.1, 0.9, 80):\n    train_f1_scores.append(f1_score(train_label.numpy().reshape(-1), (train_pred > th).numpy().reshape(-1), average=\"macro\"))\n    \nplt.plot(np.linspace(0.1, 0.9, 80), train_f1_scores, label=\"dev\")\nplt.ylim(0, 1)\n\nthreshold = np.linspace(0.1, 0.9, 80)[np.argmax(train_f1_scores)]\nprint(threshold, max(train_f1_scores))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems like the val set hasn't been used to evaluate the model during training on the entire dataset. It is a good idea to evaluate the model periodically on a separate validation set to keep track of the generalization ability of the model during training.","metadata":{}},{"cell_type":"code","source":"embeddings = emb_model(unique_feature).numpy()\n\ndel train_histories, train_label, train_labels\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code deletes the variables train_histories, train_label, and train_labels from the current namespace and frees up the memory that they were occupying by calling the garbage collector (gc.collect()).\n\nThis is done to avoid using too much memory, as these variables were used to store the training data and labels which are no longer needed for the rest of the code. By deleting them, we can free up memory and reduce the memory footprint of the program.","metadata":{}},{"cell_type":"code","source":"import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The make_env() function sets up an environment that contains an iterator called iter_test(). This iterator allows you to iterate through the test data of the competition in a streaming fashion, where the data comes in batches (called \"episodes\" in the competition) and you need to predict an action for each episode.\n\nYou can use the iter_test() iterator to generate predictions for the test data and submit them to the competition to evaluate the performance of your model.","metadata":{}},{"cell_type":"code","source":"print(unique_train_df.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This line of code should print the shape of the unique_train_df dataframe. The shape of a dataframe represents the number of rows and columns in the dataframe, and is returned as a tuple in the format (num_rows, num_columns).","metadata":{}},{"cell_type":"code","source":"# predict\neps = epsilon\nfor (sample_submission, test) in iter_test:\n    level_group = ['0-4', '5-12', '13-22'].index(test[\"level_group\"].values[0])\n    for col in feature_columns:\n        test[col + \"_ix\"] = np.vectorize(lambda x: feature_dicts[col].get(x, 0))(test[col].fillna(\"nan_value\").astype(str)).astype(np.int32)\n\n    # update embedding table\n    test_feats = test[feature_ix_columns].merge(unique_train_df,\n                                                on=feature_ix_columns,\n                                                how=\"left\")\n    unseen_test_feats = test_feats[test_feats[\"unique_ix\"].isna()].drop_duplicates(feature_ix_columns)[feature_ix_columns]\n    new_embedding = emb_model(unseen_test_feats.values).numpy()\n\n    if new_embedding.shape[0]:\n        unseen_test_feats[\"unique_ix\"] = list(range(unique_train_df.shape[0], unique_train_df.shape[0] + new_embedding.shape[0]))\n        unique_train_df = pd.concat([unique_train_df, unseen_test_feats], axis=0)\n        embeddings = np.concatenate([embeddings, new_embedding], axis=0)\n\n    # predict\n    test_feats = test[feature_ix_columns].merge(unique_train_df,\n                                                on=feature_ix_columns,\n                                                how=\"left\")\n    pred_emb = embeddings[test_feats[\"unique_ix\"].values].sum(axis=0).reshape([1, -1])\n    pred_emb /= np.maximum(eps, np.linalg.norm(pred_emb, ord=2))\n    sample_submission[\"correct\"] = tf.cast(tf.math.sigmoid(predictors[level_group](pred_emb)) > threshold, \"int32\").numpy()[0]\n\n    env.predict(sample_submission)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code is part of a for loop that iterates over test data. For each test sample, it first retrieves the corresponding feature indices and maps them to the precomputed embeddings. If any of the feature indices were not seen during training, the corresponding embedding is computed and added to the embeddings matrix. Then, the embeddings are summed and normalized, and the resulting vector is fed into a logistic regression predictor to make a binary classification (correct or incorrect) prediction. The prediction is then written to the sample_submission dataframe and submitted using the env.predict() method.","metadata":{}},{"cell_type":"code","source":"print(unique_train_df.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This will print the shape of the unique_train_df dataframe which contains unique feature indices and their corresponding embeddings.","metadata":{}},{"cell_type":"code","source":"# check prediction\n\ndf = pd.read_csv('submission.csv')\nprint( df.shape )\ndf.head(10)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Without seeing the contents of the submission.csv file, it's hard to say what's going on. However, assuming that the file has been properly formatted and contains predictions for each test sample, the shape of the dataframe should reflect the number of samples and the number of columns in the file. To check the shape of the dataframe, you can simply use the following code.\nThis will print out the number of rows and columns in the dataframe. You can then use the head method to inspect the first few rows of the dataframe and make sure the predictions are in the correct format.","metadata":{}},{"cell_type":"code","source":"print(df.correct.mean())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The mean value of the \"correct\" column in the submission file is the proportion of predicted correct answers, so it represents the accuracy of the model on the test set.","metadata":{}}]}