{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Homework 2-1 Phoneme Classification**","metadata":{"id":"OYlaRwNu7ojq"}},{"cell_type":"markdown","source":"## The DARPA TIMIT Acoustic-Phonetic Continuous Speech Corpus (TIMIT)\nThe TIMIT corpus of reading speech has been designed to provide speech data for the acquisition of acoustic-phonetic knowledge and for the development and evaluation of automatic speech recognition systems.\n\nThis homework is a multiclass classification task, \nwe are going to train a deep neural network classifier to predict the phonemes for each frame from the speech corpus TIMIT.\n\nlink: https://academictorrents.com/details/34e2b78745138186976cbc27939b1b34d18bd5b3","metadata":{"id":"emUd7uS7crTz"}},{"cell_type":"markdown","source":"## Download Data\nDownload data from google drive, then unzip it.\n\nYou should have `timit_11/train_11.npy`, `timit_11/train_label_11.npy`, and `timit_11/test_11.npy` after running this block.<br><br>\n`timit_11/`\n- `train_11.npy`: training data<br>\n- `train_label_11.npy`: training label<br>\n- `test_11.npy`:  testing data<br><br>\n\n**notes: if the google drive link is dead, you can download the data directly from Kaggle and upload it to the workspace**\n\n下载链接在这里：https://www.kaggle.com/code/tamakoyl/2021hw02phoneme/data\n\n","metadata":{"id":"KVUGfWTo7_Oj"}},{"cell_type":"code","source":"from google.colab import drive\ndrive.mount('/content/drive')","metadata":{"id":"svEw2MGC-tXh","outputId":"3263cc7f-0dbc-477e-95cd-623ea2224d50"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!gdown --id '1HPkcmQmFGu-3OknddKIa5dNDsR05lIQR' --output data.zip\n!unzip data.zip\n!ls ","metadata":{"id":"OzkiMEcC3Foq","outputId":"c9a9bb8e-afaa-48d1-e0ab-c17d4626adef"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preparing Data\nLoad the training and testing data from the `.npy` file (NumPy array).","metadata":{"id":"_L_4anls8Drv"}},{"cell_type":"code","source":"import numpy as np\n\nprint('Loading data ...')\n\ndata_root='./timit_11/'\ntrain = np.load(data_root + 'train_11.npy') # 加载训练数据集的自变量\ntrain_label = np.load(data_root + 'train_label_11.npy') #加载训练数据集的标签\ntest = np.load(data_root + 'test_11.npy') # 加载测试数据集的自变量\n\nprint('Size of training data: {}'.format(train.shape))\nprint('Size of testing data: {}'.format(test.shape))","metadata":{"id":"IJjLT8em-y9G","outputId":"cb99317b-b846-4fea-9c7f-cc7f206b3c8b"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Dataset\n这里注意要把自变量存储为float类型的numpy（把每一段语音（11 * 39）flatten为（1 * 429）的一维input）\n\n每一段语音由11个frames组成、而每一个frame包含39个类别","metadata":{"id":"us5XW_x6udZQ"}},{"cell_type":"code","source":"import torch\nfrom torch.utils.data import Dataset\n\n# 每个数据集类的定义都有3个部分\n# 初始化的init、过去数据的getitem、返回数据集长度的len\nclass TIMITDataset(Dataset):\n    # 数据集的初始化，这里要按照应变量是否存在分类讨论（因为有可能是train数据集也有可能是test数据集）\n    def __init__(self, X, y=None):\n        self.data = torch.from_numpy(X).float() # 自变量存储为float类型的numpy（把每一段语音（11*39）flatten为（1*429）的一维input）\n        if y is not None:\n            y = y.astype(np.int)\n            self.label = torch.LongTensor(y)  # 这是一个神经网络分类任务，所以应变量是一个int类型的tensor\n        else:\n            self.label = None\n\n    def __getitem__(self, idx):\n        if self.label is not None:\n            return self.data[idx], self.label[idx]\n        else:\n            return self.data[idx]\n\n    def __len__(self):\n        return len(self.data)\n","metadata":{"id":"Fjf5EcmJtf4e"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Split the labeled data into a training set and a validation set, you can modify the variable `VAL_RATIO` to change the ratio of validation data.","metadata":{"id":"otIC6WhGeh9v"}},{"cell_type":"code","source":"# VAL_RATIO为0.2，就是把train、train_label数据集的80%用于训练、剩下20%用于交叉验证\nVAL_RATIO = 0.2\n\npercent = int(train.shape[0] * (1 - VAL_RATIO)) # percent就是前80%的数据的index\n# 虽然以下这些变量都是二维数组，但是都是对行做变换，所以只有一个[]\ntrain_x, train_y, val_x, val_y = train[:percent], train_label[:percent], train[percent:], train_label[percent:]\nprint('Size of training set: {}'.format(train_x.shape))\nprint('Size of validation set: {}'.format(val_x.shape))","metadata":{"id":"sYqi_lAuvC59","outputId":"b2d10cc2-6458-418b-e822-50ed5546c0a9"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Create a data loader from the dataset, feel free to tweak the variable `BATCH_SIZE` here.\n\n这里只创建train和val的训练数据集，test的数据集等测试的步骤再说","metadata":{"id":"nbCfclUIgMTX"}},{"cell_type":"code","source":"BATCH_SIZE = 64\n\nfrom torch.utils.data import DataLoader\n\ntrain_set = TIMITDataset(train_x, train_y)\nval_set = TIMITDataset(val_x, val_y)\ntrain_loader = DataLoader(train_set, batch_size=BATCH_SIZE, shuffle=True) #only shuffle the training data\nval_loader = DataLoader(val_set, batch_size=BATCH_SIZE, shuffle=False)  # 加载数据只需要3个参数就可以了，其他默认，注意验证数据集不需要洗牌！","metadata":{"id":"RUCbQvqJurYc","outputId":"b6bb08bb-9661-4740-ff04-693b2c6ea065"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cleanup the unneeded variables to save memory.<br>\n\n**notes: if you need to use these variables later, then you may remove this block or clean up unneeded variables later<br>the data size is quite huge, so be aware of memory usage in colab**\n\n此刻由于我们的变量很大很多（保存了很多不再需要的大变量），而之后可能只需要train_loader和val_loader就行了，其他的都删掉，避免内存爆了。","metadata":{"id":"_SY7X0lUgb50"}},{"cell_type":"code","source":"import gc\n\ndel train, train_label, train_x, train_y, val_x, val_y\ngc.collect() # 检查此刻内存占用情况（值越小越好，一开始是120，删掉上述这些变量后减小为50）","metadata":{"id":"y8rzkGraeYeN","outputId":"10b1be24-452a-4565-eac3-38251009c096"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create Model","metadata":{"id":"IRqKNvNZwe3V"}},{"cell_type":"markdown","source":"Define model architecture, you are encouraged to change and experiment with the model architecture.","metadata":{"id":"FYr1ng5fh9pA"}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\n\n# 定义模型类需要2个函数，初始化函数init、以及前向传播函数forward\nclass Classifier(nn.Module):\n    # 初始化函数init的目的是定义所有的隐藏层\n    def __init__(self):\n        super(Classifier, self).__init__()\n        self.layer1 = nn.Linear(429, 1024)\n        self.layer2 = nn.Linear(1024, 512)\n        self.layer3 = nn.Linear(512, 128)\n        self.out = nn.Linear(128, 39) \n\n        self.act_fn = nn.Sigmoid()\n    # 前向函数forward的目的是实例化所有的隐藏层\n    def forward(self, x):\n        x = self.layer1(x) # 第一个隐藏层、429->1024，然后sigmoid一下\n        x = self.act_fn(x)\n\n        x = self.layer2(x) # 第二个隐藏层、1024->512，然后sigmoid一下\n        x = self.act_fn(x)\n\n        x = self.layer3(x) # 第三个隐藏层、512->128，然后sigmoid一下\n        x = self.act_fn(x)\n\n        x = self.out(x) # 输出层、128->39\n        # 这里需要说明，线性操作的过程中不止是feature的值在改变，由于“行*列\"永远不变，所以当feature改变（即列改变），那么行也在改变\n        # 比如429->1024，实质是“[1,429]->[1024/49,1024]\" 、 128->39的实质是\"[3.35,128]->[11,39]\"\n        \n        return x","metadata":{"id":"lbZrwT6Ny0XL"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training","metadata":{"id":"VRYciXZvPbYh"}},{"cell_type":"code","source":"#check device\ndef get_device():\n  return 'cuda' if torch.cuda.is_available() else 'cpu'","metadata":{"id":"y114Vmm3Ja6o"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fix random seeds for reproducibility.\n\n修改随机数，保证每一次训练结果一样，这段代码不需要背会，用的时候直接cv即可","metadata":{"id":"sEX-yjHjhGuH"}},{"cell_type":"code","source":"# fix random seed\ndef same_seeds(seed):\n    torch.manual_seed(seed)\n    if torch.cuda.is_available():\n        torch.cuda.manual_seed(seed)\n        torch.cuda.manual_seed_all(seed)  \n    np.random.seed(seed)  \n    torch.backends.cudnn.benchmark = False\n    torch.backends.cudnn.deterministic = True","metadata":{"id":"88xPiUnm0tAd"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Feel free to change the training parameters here.","metadata":{"id":"KbBcBXkSp6RA"}},{"cell_type":"code","source":"# fix random seed for reproducibility 设置相同随机数\nsame_seeds(0)\n\n# get device 获取训练装置\ndevice = get_device()\nprint(f'DEVICE: {device}')\n\n# training parameters\nnum_epoch = 20               # number of training epoch 训练20轮\nlearning_rate = 0.0001       # learning rate 这里学习率很低\n\n# the path where checkpoint saved # 设置checkpoint的存放位置，其实就是存放最优模型的位置，ckpt是文件类型、写pth也没错\nmodel_path = './model.ckpt'\n\n# create model, define a loss function, and optimizer\n# 实例化模型的3大步骤：1、实例化模型；2、实例化损失函数；3、实例化优化器\nmodel = Classifier().to(device)\ncriterion = nn.CrossEntropyLoss() # pytorch中如果采用交叉验证的损失函数，那么一定是分类任务，这时候softmax是包含在crossEntropyLoss函数中的，不需要额外调用\noptimizer = torch.optim.Adam(model.parameters(), lr=learning_rate)","metadata":{"id":"QTp3ZXg1yO9Y","outputId":"c1420782-419a-4554-eff4-b05d85cc9713"},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# start training\n\nbest_acc = 0.0\nfor epoch in range(num_epoch):\n    train_acc = 0.0\n    train_loss = 0.0\n    val_acc = 0.0\n    val_loss = 0.0\n\n    # training\n    model.train() # set the model to training mode\n    for i, data in enumerate(train_loader):\n        inputs, labels = data\n        inputs, labels = inputs.to(device), labels.to(device)\n        optimizer.zero_grad() # 每次循环之前梯度清零\n        outputs = model(inputs) \n        batch_loss = criterion(outputs, labels) # softMax包含在crossEntropyLoss函数中的，不需要额外调用\n        _, train_pred = torch.max(outputs, 1) # 选取预测结果中39个类别中，预测概率最大的类别作为预测值\n        batch_loss.backward() \n        optimizer.step() \n\n        train_acc += (train_pred.cpu() == labels.cpu()).sum().item()  # 此时的train_acc是在计算预测正确的语言的个数，如果当前语言预测正确，则train_acc+1\n        train_loss += batch_loss.item() # train_loss用于计算总的训练的loss\n\n    # validation\n    if len(val_set) > 0:\n        model.eval() # set the model to evaluation mode\n        with torch.no_grad():\n            for i, data in enumerate(val_loader):\n                inputs, labels = data\n                inputs, labels = inputs.to(device), labels.to(device)\n                outputs = model(inputs)\n                batch_loss = criterion(outputs, labels) \n                _, val_pred = torch.max(outputs, 1) \n            \n                val_acc += (val_pred.cpu() == labels.cpu()).sum().item() # 作用和train_acc一样\n                val_loss += batch_loss.item()\n\n            # 打印每一轮epoch的训练精度、训练loss、验证精度、验证loss\n            print('[{:03d}/{:03d}] Train Acc: {:3.6f} Loss: {:3.6f} | Val Acc: {:3.6f} loss: {:3.6f}'.format(\n                epoch + 1, num_epoch, train_acc/len(train_set), train_loss/len(train_loader), val_acc/len(val_set), val_loss/len(val_loader)\n            ))\n\n            # if the model improves, save a checkpoint at this epoch\n            # 这样一来最后保存的模型文件只有一个，且保存的这一个模型文件是精度最高的那一次epoch所训练的模型文件\n            if val_acc > best_acc:\n                best_acc = val_acc\n                torch.save(model.state_dict(), model_path)\n                print('saving model with acc {:.3f}'.format(best_acc/len(val_set)))\n    else: # 这个else是给验证集为0时准备的，当我们设置超参VAL_RATIO为0时，验证集为0，就要执行这个else\n        print('[{:03d}/{:03d}] Train Acc: {:3.6f} Loss: {:3.6f}'.format(\n            epoch + 1, num_epoch, train_acc/len(train_set), train_loss/len(train_loader)\n        ))\n\n# if not validating, save the last epoch\n# 这一步也是为了防止没有验证集所写的，如果验证集为0，则不会有最优模型被保存，为防止出bug，设置：当没有验证集时，选取最后一次epoch训练的模型为最优模型并保存\nif len(val_set) == 0:\n    torch.save(model.state_dict(), model_path)\n    print('saving model at last epoch')\n","metadata":{"id":"CdMWsBs7zzNs","outputId":"5d72ee3b-87c4-4fc5-acb7-6da5ae24a061"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Testing","metadata":{"id":"1Hi7jTn3PX-m"}},{"cell_type":"markdown","source":"Create a testing dataset, and load model from the saved checkpoint.","metadata":{"id":"NfUECMFCn5VG"}},{"cell_type":"code","source":"# create testing dataset\ntest_set = TIMITDataset(test, None)\ntest_loader = DataLoader(test_set, batch_size=BATCH_SIZE, shuffle=False)\n# del test_set # 删除不再需要的大型变量，防止内存爆掉\n\n# create model and load weights from checkpoint\nmodel = Classifier().to(device)\nmodel.load_state_dict(torch.load(model_path))","metadata":{"id":"1PKjtAScPWtr","outputId":"1631c4d5-465d-4e58-8223-9f247369662c"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Make prediction.","metadata":{"id":"940TtCCdoYd0"}},{"cell_type":"code","source":"predict = []\nmodel.eval() # set the model to evaluation mode\nwith torch.no_grad():\n    for i, data in enumerate(test_loader):\n        inputs = data\n        inputs = inputs.to(device)\n        outputs = model(inputs)\n        _, test_pred = torch.max(outputs, 1) # get the index of the class with the highest probability\n\n        for y in test_pred.cpu().numpy(): # 这里之所以使用for是因为，预测值不止一个（因为模型是按照batch加载的，训练也是每batch个音频一起训练，所以得到的应该是batch个预测值）\n            predict.append(y)","metadata":{"id":"84HU5GGjPqR0"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Write prediction to a CSV file.\n\nAfter finish running this block, download the file `prediction.csv` from the files section on the left-hand side and submit it to Kaggle.","metadata":{"id":"AWDf_C-omElb"}},{"cell_type":"code","source":"with open('prediction.csv', 'w') as f:\n    f.write('Id,Class\\n')\n    for i, y in enumerate(predict):\n        f.write('{},{}\\n'.format(i, y))","metadata":{"id":"GuljYSPHcZir"},"execution_count":null,"outputs":[]}]}