{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Understanding Learning Rate in Graphnet:\n\nIn this Kaggle notebook, I will explore the concept of learning rate in the context of training deep learning models. The learning rate is a critical hyperparameter that determines the step size at which the model's parameters are updated during the optimization process. Choosing the right learning rate is essential for achieving good convergence, preventing oscillations around the minimum, and ensuring efficient training. I will discuss a common strategy called \"**cyclical learning rate**\" or \"**learning rate warm-up**\" where the learning rate is first increased and then decreased. I will illustrate this concept using a PyTorch example with custom learning rate scheduler code and an example to show its impact on the training process.","metadata":{}},{"cell_type":"markdown","source":"\nI obtained the build_model function from [graphnet example](https://www.kaggle.com/code/rasmusrse/graphnet-example) and within the latter half of the function, the learning rate was established.","metadata":{}},{"cell_type":"code","source":"def build_model(config: Dict[str,Any], train_dataloader: Any) -> StandardModel:\n    \n    detector = IceCubeKaggle(\n        graph_builder=KNNGraphBuilder(nb_nearest_neighbours=8),\n    )\n    gnn = DynEdge(\n        nb_inputs=detector.nb_outputs,\n        global_pooling_schemes=[\"min\", \"max\", \"mean\", \"sum\"]\n    )\n    \n    task = DirectionReconstructionWithKappa(\n        hidden_size=gnn.nb_outputs,\n        target_labels=config[\"target\"],\n        loss_function=VonMisesFisher3DLoss(),\n    )\n    prediction_columns = [config[\"target\"] + \"_x\", \n                            config[\"target\"] + \"_y\", \n                            config[\"target\"] + \"_z\", \n                            config[\"target\"] + \"_kappa\" ]\n                            \n    additional_attributes = ['zenith', 'azimuth', 'event_id']\n\n    model = StandardModel(\n        detector=detector,\n        gnn=gnn,\n        tasks=[task],\n        optimizer_class=Adam,\n        optimizer_kwargs={\"lr\": 1e-03, \"eps\": 1e-03},\n        scheduler_class=PiecewiseLinearLR,\n        scheduler_kwargs={\n            \"milestones\": [\n                0,\n                len(train_dataloader) / 2,\n                len(train_dataloader) * config[\"fit\"][\"max_epochs\"],\n            ],\n            \"factors\": [1e-02, 1, 1e-02],\n        },\n        scheduler_config={\n            \"interval\": \"step\",\n        },\n    )\n    model.prediction_columns = prediction_columns\n    model.additional_attributes = additional_attributes\n    \n    return model","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. This code defines a function build_model that takes a configuration dictionary config and a training dataloader train_dataloader as input arguments and returns a StandardModel object.\n1. Instantiate the detector: An IceCubeKaggle detector is created, which uses a KNNGraphBuilder with 8 nearest neighbors to build the graph\n1. Instantiate the GNN: A DynEdge Graph Neural Network (GNN) model is created. The input size of the GNN is determined by the number of outputs from the detector. The model uses four global pooling schemes: \"min\", \"max\", \"mean\", and \"sum\".\n1. Define the task: A DirectionReconstructionWithKappa task is created for the model. The task takes the GNN output size as the input size for the hidden layer and uses the VonMisesFisher3DLoss as the loss function.\n1. Define the prediction columns and additional attributes: The prediction columns include the target's x, y, z components, and kappa. Additional attributes include zenith, azimuth, and event_id.\n1. Instantiate the StandardModel: The StandardModel object is created with the detector, GNN, task, optimizer class (Adam), optimizer kwargs, scheduler class (PiecewiseLinearLR), scheduler kwargs, and scheduler config.\n1. Set the prediction columns and additional attributes for the model.\n1. Return the model: The function returns the built StandardModel object.\n\n\nIn summary, this function builds a model with an IceCubeKaggle detector, DynEdge GNN, DirectionReconstructionWithKappa task, and custom learning rate scheduler (PiecewiseLinearLR). It sets the prediction columns and additional attributes and returns the constructed model.","metadata":{}},{"cell_type":"markdown","source":"And here is the PiecewiseLinearLR class from [callbacks.py](https://github.com/graphnet-team/graphnet/blob/0aadff5376130168c8760b907506aa93e56e2bd8/src/graphnet/training/callbacks.py)","metadata":{}},{"cell_type":"code","source":"class PiecewiseLinearLR(_LRScheduler):\n    \"\"\"Interpolate learning rate linearly between milestones.\"\"\"\n\n    def __init__(\n        self,\n        optimizer: Optimizer,\n        milestones: List[int],\n        factors: List[float],\n        last_epoch: int = -1,\n        verbose: bool = False,\n    ):\n        \"\"\"Construct `PiecewiseLinearLR`.\n        For each milestone, denoting a specified number of steps, a factor\n        multiplying the base learning rate is specified. For steps between two\n        milestones, the learning rate is interpolated linearly between the two\n        closest milestones. For steps before the first milestone, the factor\n        for the first milestone is used; vice versa for steps after the last\n        milestone.\n        Args:\n            optimizer: Wrapped optimizer.\n            milestones: List of step indices. Must be increasing.\n            factors: List of multiplicative factors. Must be same length as\n                `milestones`.\n            last_epoch: The index of the last epoch.\n            verbose: If ``True``, prints a message to stdout for each update.\n        \"\"\"\n        # Check(s)\n        if milestones != sorted(milestones):\n            raise ValueError(\"Milestones must be increasing\")\n        if len(milestones) != len(factors):\n            raise ValueError(\n                \"Only multiplicative factor must be specified for each milestone.\"\n            )\n\n        self.milestones = milestones\n        self.factors = factors\n        super().__init__(optimizer, last_epoch, verbose)\n\n    def _get_factor(self) -> np.ndarray:\n        # Linearly interpolate multiplicative factor between milestones.\n        return np.interp(self.last_epoch, self.milestones, self.factors)\n\n    def get_lr(self) -> List[float]:\n        \"\"\"Get effective learning rate(s) for each optimizer.\"\"\"\n        if not self._get_lr_called_within_step:\n            warnings.warn(\n                \"To get the last learning rate computed by the scheduler, \"\n                \"please use `get_last_lr()`.\",\n                UserWarning,\n            )\n\n        return [base_lr * self._get_factor() for base_lr in self.base_lrs]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The PiecewiseLinearLR class is a custom learning rate scheduler that inherits from the _LRScheduler class in PyTorch. This scheduler linearly interpolates the learning rate between specified milestones, allowing for a gradual change in the learning rate during training.\n\nHere is an explanation of the class methods:\n\n__init__: The constructor takes the following arguments:\n\noptimizer: The optimizer for which the learning rate is scheduled.\nmilestones: A list of step indices at which the learning rate should change. The list must be in increasing order.\nfactors: A list of multiplicative factors corresponding to the milestones. The length of this list must be the same as the length of milestones.\nlast_epoch: The index of the last epoch. Default value is -1.\nverbose: If True, the scheduler prints a message to stdout for each update.\nThe constructor performs some checks to ensure that the milestones are in increasing order and that the number of factors matches the number of milestones. It then initializes the base class _LRScheduler with the optimizer, last_epoch, and verbose flag.\n\n_get_factor: This private method linearly interpolates the multiplicative factor between milestones using NumPy's np.interp function. It takes the last epoch (current step) as the input value and interpolates between the milestones and their corresponding factors.\n\nget_lr: This method calculates the effective learning rates for each optimizer by multiplying the base learning rates with the interpolated factor from the _get_factor method. If the learning rate is requested without being called within a step, a warning is issued suggesting to use get_last_lr() to obtain the last computed learning rate.\n\nThis custom learning rate scheduler allows you to define a piecewise linear learning rate schedule where the learning rate changes linearly between specified milestones. This can be useful for implementing cyclical learning rate strategies or learning rate warm-up schedules.","metadata":{}},{"cell_type":"markdown","source":"### Example:\n\nLet's consider an example to understand how the learning rate changes using the PiecewiseLinearLR custom scheduler.\n\nSuppose we have the following configuration for the milestones and factors:\n\nmilestones: [0, 100, 200]\n\nfactors: [0.01, 1, 0.01]\n\nThis means we have 3 milestones at steps 0, 100, and 200, with corresponding multiplicative factors 0.01, 1, and 0.01. Let's also assume the base learning rate is 0.001.\n\nNow, let's see how the learning rate changes across different steps (epochs):\n\nStep 0: At the first milestone (0), the factor is 0.01. The learning rate will be the base learning rate (0.001) multiplied by the factor (0.01), resulting in a learning rate of 0.001 * 0.01 = 0.00001.\n\nStep 50: This step is between the first and second milestones (0 and 100). The multiplicative factor is linearly interpolated between 0.01 and 1. Since step 50 is halfway between steps 0 and 100, the interpolated factor will be 0.5 * (1 - 0.01) + 0.01 = 0.505. The learning rate will be 0.001 * 0.505 = 0.000505.\n\nStep 100: At the second milestone (100), the factor is 1. The learning rate will be the base learning rate (0.001) multiplied by the factor (1), resulting in a learning rate of 0.001 * 1 = 0.001.\n\nStep 150: This step is between the second and third milestones (100 and 200). The multiplicative factor is linearly interpolated between 1 and 0.01. Since step 150 is halfway between steps 100 and 200, the interpolated factor will be 0.5 * (0.01 - 1) + 1 = 0.505. The learning rate will be 0.001 * 0.505 = 0.000505.\n\nStep 200: At the third milestone (200), the factor is 0.01. The learning rate will be the base learning rate (0.001) multiplied by the factor (0.01), resulting in a learning rate of 0.001 * 0.01 = 0.00001.\n\nAs we can see, the learning rate changes between the specified milestones by linearly interpolating the multiplicative factors. This allows for a gradual and controlled change in the learning rate during training, which can help with optimization and convergence.","metadata":{}},{"cell_type":"markdown","source":"### Why learning rate warm-up?\n\nThe technique of first increasing and then decreasing the learning rate is known as the \"learning rate warm-up\" or \"cyclical learning rate\" strategy. The main motivation behind this approach is to improve the optimization process and achieve better model convergence.\n\nHere are a few reasons why this strategy can be helpful:\n\nImproved convergence: Increasing the learning rate initially can help the model to escape shallow local minima or saddle points in the loss landscape. This can result in the model converging to a better solution. As the learning rate is decreased later, the optimizer can then fine-tune the model's parameters, ensuring a more accurate convergence to the global minimum or a better local minimum.\n\nAdaptability: The cyclical learning rate can adapt better to the changing nature of the loss landscape. Initially, when the model's parameters are far from optimal, a larger learning rate can help make bigger updates, allowing the model to learn faster. Later on, as the model gets closer to the optimal solution, a smaller learning rate can provide more precise updates, avoiding oscillations around the minimum.\n\nRegularization effect: The cyclical learning rate can act as a form of implicit regularization. When the learning rate is high, the model's parameters are updated more aggressively, effectively introducing some noise into the optimization process. This noise can help prevent the model from overfitting to the training data, improving its generalization capabilities.\n\nFaster training: Using a warm-up or cyclical learning rate can help the model reach a good solution faster, as the optimizer can traverse the loss landscape more efficiently. This can result in reduced training time compared to using a constant learning rate throughout the training process.\n\nIn summary, the strategy of first increasing and then decreasing the learning rate can provide better convergence, adaptability, regularization, and faster training, ultimately improving the model's performance.","metadata":{}}]}