{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"![TPU](https://diamond-thumbnails.s3.us-west-2.amazonaws.com/thinkbigcms/Product/logo/6e276569-90f0-4491-a0ed-45b07b8b05eb.png?hash=f54001f5e45cdaa749504dafc8d86bc4) \n\n\nTPU is an accelerator available on **colab** and **kaggle** which provides a way to train **tensorflow** models way faster than on gpus. To train any data on TPU the dataset has to be converted into **tfrecords** format. \nIn this notebook we demonstarte:\n* how to parse tfrecords \n    * tfrecords created [with this notebook](https://www.kaggle.com/code/nazmuddhohaansary/tfrecords-for-tpu-training) \n    * the dataset is uploaded [in this dataset](https://www.kaggle.com/datasets/ocrteamriad/dl-sprint-tfrecords)  \n     \n    \n* converting raw audio features to **Log Mel Spectogram** features \n    * [in this notebook](https://www.kaggle.com/code/nazmuddhohaansary/logmelspctogram-basic-cnn-attention-modeling) under **Log Mel Spectogram** section we go step by  step of what happens when the features are extracted\n\n\n* training CNN-Attention with Positional Attention models in TPU\n\n\n### Useful links to understand tfrecords and TPU's \n\n**TPU for (~20x)faster training** \n* [what is TPU and why do we need them](https://www.quora.com/What-is-TPU-and-GPU-Why-and-when-do-we-need-them)\n* [Kaggle TPU a-z](https://www.kaggle.com/docs/tpu)\n\n**TFRecords**\n* [Official Tensorflow Doc](https://www.tensorflow.org/tutorials/load_data/tfrecord)\n* [basics](https://www.kaggle.com/code/ryanholbrook/tfrecords-basics/notebook)","metadata":{}},{"cell_type":"code","source":"!pip install image-classifiers","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* We locate the tfrecods by file patterns.We use star as wild card entry\n* While training with tfrecords **we must not load the data locally**. We have to use **GCS Buckets** to load the data. \n    * kaggle_datasets api provides a way to access both public and private GCS(google cloud storage). Here we are using public data but private datasets can also be used.\n* **PER_REPLICA_BATCH_SIZE**  global batch size while training will be **8 times the PER_REPLICA_BATCH_SIZE** we provide \n\n* **REC_SIZE=256** simply means while creating the tfrecords , we stored 256 audio files with their labels in one tfrecord","metadata":{}},{"cell_type":"code","source":"from kaggle_datasets import KaggleDatasets\nimport os \n#------------------------------\n# change able params\n#------------------------------\nTRAIN_GCS_PATTERNS      = [os.path.join(KaggleDatasets().get_gcs_path(\"dl-sprint-tfrecords\"),\"voted\",\"*/*.tfrecord\")]\n                          \nEVAL_GCS_PATTERNS       = [os.path.join(KaggleDatasets().get_gcs_path(\"dl-sprint-tfrecords\"),\"eval\",\"*/*.tfrecord\")]\n\nPER_REPLICA_BATCH_SIZE  = 64                          \nEPOCHS                  = 100               \n\n#------------------------------\n# fixed params while creating the tfrecords\n#------------------------------\nREC_SIZE=256  ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We import needed libraries here and collect the tfrecord paths that can be fed into [tf.data api](https://www.tensorflow.org/api_docs/python/tf/data/Dataset)which is the official way to use tfrecords ","metadata":{}},{"cell_type":"code","source":"#-------------------------------\n# imports\n#-------------------------------\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '3' \nimport tensorflow as tf\nimport random\nimport pandas as pd \nimport warnings\nimport matplotlib.pyplot as plt\nimport librosa\nimport numpy as np \nfrom tqdm.auto import tqdm\nfrom classification_models.tfkeras import Classifiers\ntqdm.pandas()\nwarnings.filterwarnings('ignore')\n\n#--------------------------\n# GCS Paths and tfrecords\n#-------------------------\ntrain_recs=[]\neval_recs =[]\ndef get_tfrecs(gcs_pattern):\n    file_paths = tf.io.gfile.glob(gcs_pattern)\n    random.shuffle(file_paths)\n    print(\"found \",len(file_paths), \"tfrecords\")\n    return file_paths\n\nfor gcs in TRAIN_GCS_PATTERNS:\n    print(\"Looking into gcs path:\",gcs)\n    train_recs+=get_tfrecs(gcs)\nfor gcs in EVAL_GCS_PATTERNS:\n    print(gcs)\n    eval_recs+=get_tfrecs(gcs)\n\nprint(\"Total Eval-recs:\",len(eval_recs))\nprint(\"Total Train-recs:\",len(train_recs))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Initialize TPU\n* we initialize the tpu cluster for using\n* based on number of **replicas** or devices we fix:\n    * BATCH_SIZE\n    * STEPS_PER_EPOCH\n    * and evaluation steps within an epoch (EVAL_STEPS)","metadata":{}},{"cell_type":"code","source":"#----------------------------------------------------------\n# Detect hardware, return appropriate distribution strategy\n#----------------------------------------------------------\n# TPU detection. No parameters necessary if TPU_NAME environment variable is set. On Kaggle this is always the case.\ntry:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()  \n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\n    tf.config.optimizer.set_jit(True)\nelse:\n    strategy = tf.distribute.get_strategy() \n    # default distribution strategy in Tensorflow. Works on CPU and single GPU.\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)\n\n#-------------------------------------\n# batching , strategy and steps\n#-------------------------------------\nif strategy.num_replicas_in_sync==1:\n    BATCH_SIZE = PER_REPLICA_BATCH_SIZE\nelse:\n    BATCH_SIZE = PER_REPLICA_BATCH_SIZE*strategy.num_replicas_in_sync\n\n# set    \nSTEPS_PER_EPOCH = (len(train_recs)*REC_SIZE)//(BATCH_SIZE)\nEVAL_STEPS      = (len(eval_recs)*REC_SIZE)//(2*BATCH_SIZE)\nprint(\"Batch Size:\",BATCH_SIZE)\nprint(\"Steps:\",STEPS_PER_EPOCH)\nprint(\"Eval Steps:\",EVAL_STEPS)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Loader : tf.data.Dataset api\n**Log Mel Spectogram**: [notebook](https://www.kaggle.com/code/nazmuddhohaansary/logmelspctogram-basic-cnn-attention-modeling) ","metadata":{}},{"cell_type":"code","source":"class data_cfg:\n    max_label_len=250\n    shuffle_buffer=2048\n    batch_size=BATCH_SIZE\n    data_shape=(64,1088,1)\n    label_shape=(250,)\n    \nclass LogMelSpectrogramProcesor(object):\n    def __init__(self,\n                 sample_rate=16000,\n                 preemphasis_coeff=0.97,\n                 frame_ms=25,\n                 stride_ms=10,\n                 center=True,\n                 num_feature_bins=64,\n                 max_width=1088):\n        \"\"\"\n            class to process and extract log mel spectrogram features from audio\n        \"\"\"\n        \n        self.sample_rate          =  sample_rate\n        self.frame_length         =  int(self.sample_rate * (frame_ms / 1000))\n        self.frame_step           =  int(self.sample_rate * (stride_ms/ 1000))\n        self.preemphasis_coeff    =  preemphasis_coeff\n        self.center               =  center\n        self.num_feature_bins     =  num_feature_bins\n        self.nfft                 =  2 ** (self.frame_length - 1).bit_length()\n        self.max_audio_len        =  self.frame_step*max_width -1  \n        \n    #------------------------------------------------------------\n    # generally useable functions for audio processing\n    #------------------------------------------------------------\n    def load_data(self,path):\n        \"\"\"loads a wav\"\"\"\n        wave,_= librosa.load(path, sr=self.sample_rate, mono=True)\n        wave=np.trim_zeros(wave)\n        return tf.cast(wave,tf.float32)\n        \n    def normalize_signal(self,signal):\n        \"\"\"Normailize signal to [-1, 1] range\"\"\"\n        gain = 1.0 / (tf.reduce_max(tf.abs(signal), axis=-1) + 1e-9)\n        return signal * gain\n    \n    def normalize_audio_feature(self,audio_feature):\n        \"\"\"Mean and variance normalization\"\"\"\n        axis=None\n        mean = tf.reduce_mean(audio_feature, axis=axis, keepdims=True)\n        std_dev = tf.math.sqrt(tf.math.reduce_variance(audio_feature, axis=axis, keepdims=True) + 1e-9)\n        return (audio_feature - mean) / std_dev\n        \n    def preemphasis(self,signal):\n        \"\"\"\n        Apply Pre-emphasis with the defined preemphasis coefficient \n        ** preemphasis: the intentional alteration of the relative strengths \n        of signals at different frequencies (as in radio and in disc recording) \n        to reduce adverse effects (as noise) in the following parts of the system.\n\n        \"\"\"\n        s0 = tf.expand_dims(signal[0], axis=-1)\n        s1 = signal[1:] - self.preemphasis_coeff * signal[:-1]\n        return tf.concat([s0, s1], axis=-1)\n    \n    def pad_signal(self,signal):\n        '''\n            pads a 1d array with 0\n        '''\n        pad_size = self.max_audio_len - tf.shape(signal)[0]\n        paddings = [[0,pad_size]] # assign here, during graph execution\n        return tf.pad(signal, paddings)\n    #------------------------------------------------------------\n    # log mel spectrogram specific\n    #------------------------------------------------------------\n    def log10(self,x):\n        numerator = tf.math.log(x)\n        denominator = tf.math.log(tf.constant(10, dtype=numerator.dtype))\n        return numerator / denominator\n    def power_to_db(self,S,amin=1e-10,top_db=80.0):\n        log_spec = 10.0 * self.log10(tf.maximum(amin, S))\n        log_spec -= 10.0 * self.log10(tf.maximum(amin, 1.0))\n        log_spec = tf.maximum(log_spec, tf.reduce_max(log_spec) - top_db)\n        return log_spec\n    \n    \n    def stft(self,signal):\n        if self.center:\n            signal = tf.pad(signal, [[self.nfft // 2, self.nfft // 2]], mode=\"REFLECT\")\n        window = tf.signal.hann_window(self.frame_length, periodic=True)\n        left_pad = (self.nfft - self.frame_length) // 2\n        right_pad = self.nfft - self.frame_length - left_pad\n        window = tf.pad(window, [[left_pad, right_pad]])\n        framed_signals = tf.signal.frame(signal, frame_length=self.nfft, frame_step=self.frame_step)\n        framed_signals *= window\n        return tf.square(tf.abs(tf.signal.rfft(framed_signals, [self.nfft])))\n        \n    def compute_log_mel_spectrogram(self,signal):\n        spectrogram = self.stft(signal)\n        linear_to_weight_matrix = tf.signal.linear_to_mel_weight_matrix(num_mel_bins=self.num_feature_bins,\n                                                                        num_spectrogram_bins=spectrogram.shape[-1],\n                                                                        sample_rate=self.sample_rate,\n                                                                        lower_edge_hertz=0.0,\n                                                                        upper_edge_hertz=(self.sample_rate / 2),)\n        mel_spectrogram = tf.tensordot(spectrogram, linear_to_weight_matrix, 1)\n        return self.power_to_db(mel_spectrogram)\n    #------------------------------------------------------------\n    def extract(self,signal):\n        \"\"\"\n            for parsing tfrecords\n        \"\"\"\n        # normalize\n        signal = self.normalize_signal(signal)\n        # preemphasis\n        signal = self.preemphasis(signal)\n        # pad data\n        signal = self.pad_signal(signal)\n        # log mel spectrogram features\n        features = self.compute_log_mel_spectrogram(signal)\n        features = tf.transpose(features)\n        # expand axis\n        features = tf.expand_dims(features, axis=-1)\n        # normalize feats\n        features = self.normalize_audio_feature(features)\n        return features\n\n    def __call__(self,path):\n        \"\"\"\n        Extract speech features from signals \n        * normalizes audio signal\n        * preemphasis on audio\n        * computes log mel spectrogram\n        * normalizes the features\n        \"\"\"\n        # signal\n        signal=self.load_data(path)\n        return self.extract(signal)\n        \n#------------------------------------------------------- \nlms=LogMelSpectrogramProcesor()\n\n#------------------------------\n# parsing tfrecords log mel spectogram\n#------------------------------\ndef pad_label(label):\n    '''\n        pads a 1d array with 0\n    '''\n    pad_size = data_cfg.max_label_len - tf.shape(label)[0]\n    paddings = [[0,pad_size]] # assign here, during graph execution\n    return tf.pad(label, paddings)\n\ndef read_raw_audio(audio):\n    wave,rate = tf.audio.decode_wav(audio, desired_channels=1, desired_samples=-1)\n    return tf.reshape(wave, shape=[-1]) \n    \ndef preprocess_example(audio,label):\n    with tf.device(\"/CPU:0\"):\n        signal = read_raw_audio(audio)\n        label = tf.strings.to_number(tf.strings.split(label), out_type=tf.int32)\n        return signal,label\n\ndef data_input_fn(recs): \n    '''\n      This Function generates data from gcs\n      * The parser function should look similiar now because of datasetEDA\n    '''\n    def _parser(example):   \n        feature ={  'audio' : tf.io.FixedLenFeature([],tf.string) ,\n                    'label' : tf.io.FixedLenFeature([],tf.string) \n        }    \n        example=tf.io.parse_single_example(example,feature)\n        audio,label=preprocess_example(**example)\n        data=lms.extract(audio)\n        data=tf.reshape(data,data_cfg.data_shape)\n        label = pad_label(label)\n        label=tf.reshape(label,data_cfg.label_shape)\n        \n        \n        return data,label\n    # fixed code (for almost all tfrec training)\n    dataset = tf.data.TFRecordDataset(recs)\n    dataset = dataset.map(_parser)\n    dataset = dataset.shuffle(data_cfg.shuffle_buffer,reshuffle_each_iteration=True)\n    dataset = dataset.repeat()\n    dataset = dataset.batch(data_cfg.batch_size,drop_remainder=True)\n    dataset = dataset.prefetch(tf.data.experimental.AUTOTUNE)\n    dataset = dataset.apply(tf.data.experimental.ignore_errors())\n    return dataset","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ds=data_input_fn(train_recs)\neval_ds =data_input_fn(eval_recs)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#------------------------------\n# view data\n#------------------------------\n    \nfor x,y in eval_ds.take(1):\n    plt.figure(figsize=(20,20))\n    plt.imshow(x[0])\n    plt.show()\n    print(x.shape)\n    print(\"label:\",y[0])\n    print(y.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling","metadata":{}},{"cell_type":"code","source":"#-----------------------------------\n#for creating Embedding Weights\n#-----------------------------------\nimport torch\nimport torch.nn as nn\n#--------------------------------------------------------------------\n# attention modules\n#--------------------------------------------------------------------\n\nclass DotAttention(tf.keras.layers.Layer):\n    '''\n            Calculate the attention weights.\n\n            args:\n                q   : query shape == (..., seq_len_q, depth)\n                k   : key shape == (..., seq_len_k, depth)\n                v   : value shape == (..., seq_len_v, depth_v)\n                mask: Float tensor with shape broadcastable to (..., seq_len_q, seq_len_k). Defaults to None.\n            returns:\n                output, attention_weights\n            NOTES:\n            * q, k, v must have matching leading dimensions.\n            * k, v must have matching penultimate dimension, i.e.: seq_len_k = seq_len_v.\n            * The mask has different shapes depending on its type(padding or look ahead) but it must be broadcastable for addition.\n\n    '''\n    def __init__(self):\n        super().__init__()\n        self.inf_val=-1e9\n        \n    def call(self,q, k, v, mask):\n        \n        matmul_qk = tf.matmul(q, k, transpose_b=True)  # (..., seq_len_q, seq_len_k)\n       \n        # scale matmul_qk\n        dk = tf.cast(tf.shape(k)[-1], tf.float32)\n        scaled_attention_logits = matmul_qk / tf.math.sqrt(dk)\n\n        # add the mask to the scaled tensor.\n        if mask is not None:\n            scaled_attention_logits += (mask * self.inf_val)\n\n        # softmax is normalized on the last axis (seq_len_k) so that the scores\n        # add up to 1.\n        attention_weights = tf.nn.softmax(scaled_attention_logits, axis=-1)  # (..., seq_len_q, seq_len_k)\n\n        output = tf.matmul(attention_weights, v)  # (..., seq_len_q, depth_v)\n\n        return output\n    \n#--------------------------------------------------------------------\n\nclass PositionalEncoding(tf.keras.layers.Layer):\n    '''\n    tensorflow wrapper for positional encoding layer\n    args:\n      num_seq  :   incoming sequence length\n      projection_dim  :   projection_dim\n      use_torch_weights : torch weights help converge faster for basic dot attention\n    '''\n    def __init__(self,num_seq,projection_dim,use_torch_weights=False):\n        super(PositionalEncoding, self).__init__()\n        self.use_torch_weights=use_torch_weights\n        self.projection_dim=projection_dim\n        self.num_seq = num_seq\n        self.projection = tf.keras.layers.Dense(units=projection_dim)\n        if use_torch_weights:\n            pos_emb              = nn.Embedding(num_seq+1,projection_dim)\n            pos_emb_weight       = pos_emb.weight.data.numpy()\n            self.position_embedding = tf.keras.layers.Embedding(input_dim=num_seq+1, output_dim=projection_dim,weights=[pos_emb_weight])\n        else:\n            \n            self.position_embedding = tf.keras.layers.Embedding(input_dim=num_seq, output_dim=projection_dim)\n\n    def call(self, x):\n        positions = tf.range(start=0, limit=self.num_seq, delta=1)\n        if x is None:\n            return self.position_embedding(positions)\n        encoded = self.projection(x) + self.position_embedding(positions)\n        return encoded\n    \n    def get_config(self):\n        config = super().get_config().copy()\n        config.update({'num_seq': self.num_seq,'projection_dim':self.projection_dim,\"use_torch_weights\":use_torch_weights})\n        return config\n\n#--------------------------------------------------------------------\n \nclass TransformerBlock(tf.keras.layers.Layer):\n    '''\n        transformer encoder block based on multihead self-attention\n    '''\n    def __init__(self, embed_dim, num_heads, ff_dim, rate=0.1):\n        super(TransformerBlock, self).__init__()\n        self.embed_dim=embed_dim\n        self.num_heads=num_heads\n        self.ff_dim   =ff_dim\n\n        self.att = tf.keras.layers.MultiHeadAttention(num_heads=num_heads, key_dim=embed_dim)\n        self.ffn = tf.keras.Sequential(\n            [tf.keras.layers.Dense(ff_dim, activation=\"relu\"), tf.keras.layers.Dense(embed_dim),]\n        )\n        self.layernorm1 = tf.keras.layers.LayerNormalization(epsilon=1e-6)\n        self.layernorm2 = tf.keras.layers.LayerNormalization(epsilon=1e-6)\n        self.dropout1 = tf.keras.layers.Dropout(rate)\n        self.dropout2 = tf.keras.layers.Dropout(rate)\n\n    def call(self, inputs, training):\n        attn_output = self.att(inputs, inputs)\n        attn_output = self.dropout1(attn_output, training=training)\n        out1 = self.layernorm1(inputs + attn_output)\n        ffn_output = self.ffn(out1)\n        ffn_output = self.dropout2(ffn_output, training=training)\n        return self.layernorm2(out1 + ffn_output)\n    def get_config(self):\n        config = super().get_config().copy()\n        config.update({'embed_dim': self.embed_dim,\n                       'num_heads': self.num_heads,\n                       'ff_dim':self.ff_dim})\n        return config\n    ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def attend(x,num_heads,num_blocks,reshape=True):\n    '''\n        basic self-attention wrapper\n        args:\n            x : input tensor\n            num_heads : heads to use in multihead attention\n            num_blocks: how many attention blocks to use \n    '''\n    bs,h,w,nc=x.shape\n    x = tf.keras.layers.Reshape((h*w,nc))(x)\n    x = PositionalEncoding(h*w,nc)(x)\n    for _ in range(num_blocks):\n        x=TransformerBlock(embed_dim=nc, num_heads=num_heads,ff_dim=4*nc)(x)\n    if reshape:\n        x = tf.keras.layers.Reshape((h,w,nc))(x)\n    return x\n\n\ndef create_basic_model(cfg):\n    '''\n        creates a basic cnn-seftattention-positional-attention based model\n        **flow**\n        cnn_feat(input_image)---> h//f,w//f,c shaped tensor = feat\n        attend(feat)---> self attention applied on the features= enc\n        pos_attention(enc)--> align encoded features with positional data=logits\n    '''\n    #-----------cnn feature extractor------------------\n    cnn,_ = Classifiers.get(cfg.backbone)\n    cnn = cnn(cfg.img_dim,weights=None,include_top=False)\n    inp = cnn.input\n    x   = cnn.output\n    bs,h,w,fc=x.shape\n    if fc!=cfg.embed_dim:\n        x=tf.keras.layers.Conv2D(cfg.embed_dim,3,padding='same')(x)\n    print(\"model input:\",inp.shape)\n    print(\"cnn feat:\",x.shape)\n    #-----------self attention------------------\n    x=attend(x,cfg.num_blocks,cfg.num_heads,reshape=False)\n    print(\"feat attention (seq,embed_dim):\",x.shape)\n    #-----------positional attention------------------\n    pos=PositionalEncoding(cfg.pos_max,cfg.embed_dim,use_torch_weights=True)(None)\n    attn=DotAttention()(pos,x,x,None)\n    print(\"positional attention (pos_max,embed_dim):\",attn.shape)\n    x=tf.keras.layers.Dense(cfg.logits_len)(attn)\n    print(\"logits(pos_max,logits_len):\",x.shape)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    return model\n    \n\n    \n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# CNN backbone model selection\nThe following models are available in this script **but any custom cnn backbone can be used according to input/output downsample ratio**\n* [source](https://github.com/qubvel/classification_models/blob/master/README.md)\n* The scores are for ```imagenet``` image classification challange\n\n| Model           |Acc@1|Acc@5|Time*|Source|\n|-----------------|:---:|:---:|:---:|------|\n|vgg16            |70.79|89.74|24.95|[keras](https://github.com/keras-team/keras-applications)|\n|vgg19            |70.89|89.69|24.95|[keras](https://github.com/keras-team/keras-applications)|\n|resnet18         |68.24|88.49|16.07|[mxnet](https://github.com/Microsoft/MMdnn)|\n|resnet34         |72.17|90.74|17.37|[mxnet](https://github.com/Microsoft/MMdnn)|\n|resnet50         |74.81|92.38|22.62|[mxnet](https://github.com/Microsoft/MMdnn)|\n|resnet101        |76.58|93.10|33.03|[mxnet](https://github.com/Microsoft/MMdnn)|\n|resnet152        |76.66|93.08|42.37|[mxnet](https://github.com/Microsoft/MMdnn)|\n|resnet50v2       |69.73|89.31|19.56|[keras](https://github.com/keras-team/keras-applications)|\n|resnet101v2      |71.93|90.41|28.80|[keras](https://github.com/keras-team/keras-applications)|\n|resnet152v2      |72.29|90.61|41.09|[keras](https://github.com/keras-team/keras-applications)|\n|resnext50        |77.36|93.48|37.57|[keras](https://github.com/keras-team/keras-applications)|\n|resnext101       |78.48|94.00|60.07|[keras](https://github.com/keras-team/keras-applications)|\n|densenet121      |74.67|92.04|27.66|[keras](https://github.com/keras-team/keras-applications)|\n|densenet169      |75.85|92.93|33.71|[keras](https://github.com/keras-team/keras-applications)|\n|densenet201      |77.13|93.43|42.40|[keras](https://github.com/keras-team/keras-applications)|\n|inceptionv3      |77.55|93.48|38.94|[keras](https://github.com/keras-team/keras-applications)|\n|xception         |78.87|94.20|42.18|[keras](https://github.com/keras-team/keras-applications)|\n|inceptionresnetv2|80.03|94.89|54.77|[keras](https://github.com/keras-team/keras-applications)|\n|seresnet18       |69.41|88.84|20.19|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|seresnet34       |72.60|90.91|22.20|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|seresnet50       |76.44|93.02|23.64|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|seresnet101      |77.92|94.00|32.55|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|seresnet152      |78.34|94.08|47.88|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|seresnext50      |78.74|94.30|38.29|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|seresnext101     |79.88|94.87|62.80|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|senet154         |81.06|95.24|137.36|[pytorch](https://github.com/Cadene/pretrained-models.pytorch)|\n|nasnetlarge      |**82.12**|**95.72**|116.53|[keras](https://github.com/keras-team/keras-applications)|\n|nasnetmobile     |74.04|91.54|27.73|[keras](https://github.com/keras-team/keras-applications)|\n|mobilenet        |70.36|89.39|15.50|[keras](https://github.com/keras-team/keras-applications)|\n|mobilenetv2      |71.63|90.35|18.31|[keras](https://github.com/keras-team/keras-applications)|","metadata":{}},{"cell_type":"code","source":"class cfg:\n    backbone='resnet18'\n    img_dim =(64,1088,1)     #features.shape\n    pos_max =250             #np.array(encode_label(sen)).shape\n    num_heads=8              # number of self-attention heads\n    num_blocks=4             # number of transformer blocks\n    embed_dim =256           # reduced channel for sequencing\n    logits_len =85           # len(vocab)+1\n \n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n    model=create_basic_model(cfg)\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#------------------metrics---------------------------\ndef C_acc(y_true, y_pred):\n    '''\n        calculates how many charecters are predicted correctly\n    '''\n    pad_value=0 # label pad index\n    accuracies = tf.equal(tf.cast(y_true,tf.int64), tf.argmax(y_pred, axis=2))\n    mask = tf.math.logical_not(tf.math.equal(y_true,pad_value))\n    accuracies = tf.math.logical_and(mask, accuracies)\n    accuracies = tf.cast(accuracies, dtype=tf.float32)\n    mask = tf.cast(mask, dtype=tf.float32)\n    return tf.reduce_sum(accuracies)/tf.reduce_sum(mask)\n#------------------loss--------------------------\nclass CharLoss(tf.keras.losses.Loss):\n    \"\"\"\n        a loss function to estimate charecter accuracy loss ignoring pad value\n    \"\"\"\n    def __init__(self,pad_value):\n        super(CharLoss, self).__init__(name=\"char_loss\")\n        self.pad_value=pad_value\n        self.loss_object = tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True, reduction='none')\n    def call(self, y_true, y_pred):\n        mask = tf.math.logical_not(tf.math.equal(y_true, self.pad_value))\n        loss_ = self.loss_object(y_true, y_pred)\n        mask = tf.cast(mask, dtype=loss_.dtype)\n        loss_ *= mask\n        return tf.reduce_sum(loss_)/tf.reduce_sum(mask)\n\nclass CTCLoss(tf.keras.losses.Loss):\n    \"\"\" A class that wraps the function of tf.nn.ctc_loss. \n    \n    Attributes:\n        logits_time_major: If False (default) , shape is [batch, time, logits], \n            If True, logits is shaped [time, batch, logits]. \n        blank_index: Set the class index to use for the blank label. default is\n            -1 (num_classes - 1). \n    \"\"\"\n\n    def __init__(self, logits_time_major=False, name='ctc_loss'):\n        super().__init__(name=name)\n        self.logits_time_major = logits_time_major\n\n    def call(self, y_true, y_pred):\n        \"\"\" \n            Computes CTC (Connectionist Temporal Classification) loss. \n        \"\"\"\n        y_true = tf.cast(y_true, tf.int32)\n        logit_length = tf.fill([tf.shape(y_pred)[0]], tf.shape(y_pred)[1])\n        label_length = tf.fill([tf.shape(y_true)[0]], tf.shape(y_true)[1])\n        loss = tf.nn.ctc_loss(\n            labels=y_true,\n            logits=y_pred,\n            label_length=label_length,\n            logit_length=logit_length,\n            logits_time_major=self.logits_time_major,\n            blank_index=0)\n        return tf.math.reduce_mean(loss)\n    \n# early stopping\nearly_stopping = tf.keras.callbacks.EarlyStopping(patience=5, \n                                                  verbose=1, \n                                                  mode = 'auto') \ncallbacks = [tf.keras.callbacks.ModelCheckpoint(\"model.h5\",\n                                                save_best_only=True,\n                                                save_weights_only=True,\n                                                verbose=1),\n             early_stopping]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with strategy.scope():\n\n    lr_schedule = tf.keras.experimental.CosineDecay(initial_learning_rate=0.0001,\n                                                             decay_steps=600000,\n                                                             alpha= 0.01)\n\n    model.compile(optimizer=tf.keras.optimizers.Adam(lr_schedule),\n                  loss=CharLoss(0),\n                  metrics=[C_acc])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EPOCHS=100\nhistory=model.fit(train_ds,\n                  epochs=EPOCHS,\n                  steps_per_epoch=STEPS_PER_EPOCH,\n                  verbose=1,\n                  validation_data=eval_ds,\n                  validation_steps=EVAL_STEPS, \n                  callbacks=callbacks)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**curves are produced after completing training only**","metadata":{}},{"cell_type":"code","source":"curves={}\nfor key in history.history.keys():\n    curves[key]=history.history[key]\ncurves=pd.DataFrame(curves)\ncurves.to_csv(f\"history.csv\",index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference Codeblock\n* **train as needed**\n* this inference code conveys the idea only - \n\n> its better to use the saved weights and infer on the model separately with GPU","metadata":{}},{"cell_type":"code","source":"sub=pd.read_csv(\"../input/dlsprint/sample_submission.csv\")\nsub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vocab=[ 'pad','start','end','\\u200d',\n        ' ','!',\"'\",',','-','.',':',';','=','?','।',\n        'ঁ','ং','ঃ',\n        'অ','আ','ই','ঈ','উ','ঊ','ঋ','এ','ঐ','ও','ঔ',\n        'ক','খ','গ','ঘ','ঙ',\n        'চ','ছ','জ','ঝ','ঞ',\n        'ট','ঠ','ড','ঢ','ণ',\n        'ত','থ','দ','ধ','ন',\n        'প','ফ','ব','ভ','ম',\n        'য','র','ল',\n        'শ','ষ','স','হ',\n        'া','ি','ী','ু','ূ','ৃ','ে','ৈ','ো','ৌ','্',\n        'ৎ','ড়','ঢ়','য়',\n        '০','১','২','৩','৪','৫','৬','৭','৮','৯']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TEST_WAVS=\"../input/test-wav-files-dl-sprint/test_files_wav\"\nPREDS=[]\nBATCH_SIZE=8\nfor idx in tqdm(range(0,len(sub),BATCH_SIZE)):\n    batch=[]\n    for bi in range(idx,idx+BATCH_SIZE):\n        _path=sub.iloc[bi,0]\n        _path=os.path.join(TEST_WAVS,_path).replace(\".mp3\",\".wav\")\n        signal=lms(_path)\n        batch.append(tf.expand_dims(signal,axis=0))\n    batch=tf.concat(batch,axis=0)\n    preds=model(batch,training=False)\n    for pred in preds:\n        out=np.argmax(pred,axis=-1)\n        text=[vocab[i] for i in out]\n        text=\"\".join(text)\n        PREDS.append(text)\n    ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub[\"sentence\"]=PREDS\nsub","metadata":{},"execution_count":null,"outputs":[]}]}