{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# An simple introduction to tile configuration acceleration with Pytorch","metadata":{}},{"cell_type":"markdown","source":"## Some background of the AI compiler \n**Layout configuration** in AI compilers refers to the arrangement of data in physical memory(GPUs or TPUs). This can have a significant impact on the performance of AI models, as it affects how efficiently data can be accessed by the hardware.\n\n**Tile configuration** in AI compilers refers to the way that data is divided into smaller chunks for processing. This can also have a significant impact on performance, as it can help to reduce memory bandwidth requirements and improve cache utilization.\n\nHere are some PyTorch examples of layout configuration and tile configuration:\n\n**Layout configuration**\n\nThe \n* **Channel-first layout:** This is the default layout for PyTorch tensors, and it is the most efficient layout for most GPUs. In channel-first layout, the channels of a tensor are stored contiguously in memory.\n* **Channel-last layout:** This layout is more efficient for some operations, such as convolutions on depthwise separable filters. In channel-last layout, the channels of a tensor are stored last in memory.\n* **NCHW layout:** This layout is commonly used for training and inference on CPUs. In NCHW layout, the tensors are stored in the following order: batch size, channels, height, width.\n* **NHWC layout:** This layout is commonly used for training and inference on mobile devices. In NHWC layout, the tensors are stored in the following order: batch size, height, width, channels.\n\n**Tile configuration**\n\n* **Spatial tiling:** This is the most common type of tiling, and it involves dividing the spatial dimensions of a tensor into smaller chunks. Spatial tiling can help to reduce memory bandwidth requirements and improve cache utilization.\n* **Channel tiling:** This type of tiling involves dividing the channel dimension of a tensor into smaller chunks. Channel tiling can be useful for improving the performance of convolutions on depthwise separable filters.\n* **Micro-tiling:** This type of tiling involves dividing the tensor into very small chunks. Micro-tiling can be useful for improving the performance of some operations, such as matrix multiplication.\n\nHere is an example of how to use tile configuration in PyTorch:\n","metadata":{}},{"cell_type":"code","source":"import torch, gc\ndevice = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:46:02.546855Z","iopub.execute_input":"2023-10-09T15:46:02.547279Z","iopub.status.idle":"2023-10-09T15:46:05.523803Z","shell.execute_reply.started":"2023-10-09T15:46:02.547241Z","shell.execute_reply":"2023-10-09T15:46:05.522856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\n\nclass CNN(torch.nn.Module):\n    def __init__(self):\n        super().__init__()\n\n        # Define the CNN layers\n        self.conv1 = torch.nn.Conv2d(3, 32, kernel_size=3, stride=1, padding=1)\n        self.pool1 = torch.nn.MaxPool2d(2, 2)\n        self.conv2 = torch.nn.Conv2d(32, 64, kernel_size=3, stride=1, padding=1)\n        self.pool2 = torch.nn.MaxPool2d(2, 2)\n        self.fc1 = torch.nn.Linear(64 * 7 * 7, 10)\n\n    def forward(self, x):\n        # Perform the forward pass of the CNN\n        x = self.conv1(x)\n        x = self.pool1(x)\n        x = self.conv2(x)\n        x = self.pool2(x)\n        x = x.view(-1, 64 * 7 * 7)\n        x = self.fc1(x)\n\n        return x\n\n# Create a CNN model\nmodel = CNN().to(device)\n\n# Tile the input tensor across the spatial dimensions\ninput_tensor = torch.randn(100, 3, 224, 224).to(device, dtype=torch.float32)\ntiled_input_tensor = torch.tile(input_tensor, (1, 1, 2, 2)).to(device)\n# 1.5192","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:46:05.525613Z","iopub.execute_input":"2023-10-09T15:46:05.526357Z","iopub.status.idle":"2023-10-09T15:46:08.607571Z","shell.execute_reply.started":"2023-10-09T15:46:05.526325Z","shell.execute_reply":"2023-10-09T15:46:08.606617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Perform the forward pass of the CNN on the tiled tensor\noutput_tensor = model(input_tensor).to(device)\n\n# Print the output tensor\nprint(output_tensor)\n# 3.774","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:46:08.608857Z","iopub.execute_input":"2023-10-09T15:46:08.609437Z","iopub.status.idle":"2023-10-09T15:46:13.017149Z","shell.execute_reply.started":"2023-10-09T15:46:08.609404Z","shell.execute_reply":"2023-10-09T15:46:13.016247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Perform the forward pass of the CNN on the tiled tensor\noutput_tensor = model(tiled_input_tensor)\n\n# Print the output tensor\nprint(output_tensor)\n# 10.362","metadata":{"execution":{"iopub.status.busy":"2023-10-09T15:46:13.019233Z","iopub.execute_input":"2023-10-09T15:46:13.019781Z","iopub.status.idle":"2023-10-09T15:46:13.104358Z","shell.execute_reply.started":"2023-10-09T15:46:13.019751Z","shell.execute_reply":"2023-10-09T15:46:13.103435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Performance Comparison\nThe tiled operation has almost **10 times faster at least** while almost three times memory consumption **on my local GPU for inference task**. ","metadata":{}}]}