{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nImport Libraries\n</div>","metadata":{}},{"cell_type":"code","source":"from keras.layers import (BatchNormalization,Flatten,Convolution1D,Convolution2D,Activation,Input,Dense,LSTM,GRU,GaussianDropout,Reshape,Concatenate)\nfrom tsfresh.feature_extraction import feature_calculators\nfrom keras.callbacks import ModelCheckpoint,EarlyStopping\nfrom keras.utils import Sequence, to_categorical\nfrom sklearn.metrics import mean_absolute_error\nfrom keras.callbacks import ReduceLROnPlateau\nfrom keras import losses, models, optimizers\nfrom sklearn.model_selection import KFold\nfrom tqdm import tqdm_notebook as tqdm\nfrom joblib import Parallel, delayed#Parallel对象会创建一个进程池，以便在多进程中执行每一个列表项。函数delayed是一个创建元组(function, args, kwargs)的简单技巧。\nfrom sklearn import preprocessing\nimport matplotlib.pyplot as plt\nfrom keras import backend as K\nimport tensorflow as tf\nimport lightgbm as lgb\nimport seaborn as sns\nimport random as rn\nimport pandas as pd\nimport numpy as np\nimport scipy as sp\nimport itertools\nimport warnings\nimport librosa\nimport pywt\nimport os\nimport gc\n\n\nwarnings.simplefilter(action='ignore', category=FutureWarning)\npd.options.mode.chained_assignment = None\n%matplotlib inline","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-17T10:34:14.001379Z","iopub.execute_input":"2023-03-17T10:34:14.001704Z","iopub.status.idle":"2023-03-17T10:34:14.009299Z","shell.execute_reply.started":"2023-03-17T10:34:14.001641Z","shell.execute_reply":"2023-03-17T10:34:14.008391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(tf.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-03-14T10:10:39.38614Z","iopub.execute_input":"2023-03-14T10:10:39.38658Z","iopub.status.idle":"2023-03-14T10:10:39.392005Z","shell.execute_reply.started":"2023-03-14T10:10:39.386505Z","shell.execute_reply":"2023-03-14T10:10:39.391152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nDescription\n</div>\nThe LANL Earthquake Prediction competition (https://www.kaggle.com/c/LANL-Earthquake-Prediction/overview) requires competitors to predict the time remaining (Time to failure, or TTF) until a laboratory earthquake occurs from real-time seismic data. We are given 150,000 data points of seismic data, which corresponds to 0.0375 seconds of seismic data (ordered in time).\n\nThe first place solution to LANL Earthquake Prediction is a geometric mean of a neural network (NN) solution and LightGBM (LGBM) solution. Since the NN and LGBM algorithms are very different, they each capture different parts of the signal, and blending the two together increases generalization.\n\nTo make predictions, we divide our training data into 150,000-length segments. Instead of using all of the training data, we decided to use only segments from the earthquake cycles that had exhibited higher TTF. This caused the TTF predictions to be biased higher.\n\nThe raw acoustic data itself is noisy; therefore, we utilize various packages to denoise the signal. Additionally, we inject random noise to every segment and remove the median of the segment, because we noticed the mean & median values were increasing as the laboratory experiment went forward in time. This improves generalization.\n\nAdditional solution details can be found on Kaggle's Discussion forum at https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/94390.","metadata":{}},{"cell_type":"markdown","source":"## Read in data","metadata":{}},{"cell_type":"code","source":"# raw train data import\nraw = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32}) ","metadata":{"execution":{"iopub.status.busy":"2023-03-17T09:53:08.06252Z","iopub.execute_input":"2023-03-17T09:53:08.062863Z","iopub.status.idle":"2023-03-17T09:57:50.27748Z","shell.execute_reply.started":"2023-03-17T09:53:08.062805Z","shell.execute_reply":"2023-03-17T09:57:50.276563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"raw","metadata":{"execution":{"iopub.status.busy":"2023-03-14T09:04:27.17394Z","iopub.execute_input":"2023-03-14T09:04:27.174272Z","iopub.status.idle":"2023-03-14T09:04:27.213125Z","shell.execute_reply.started":"2023-03-14T09:04:27.174212Z","shell.execute_reply":"2023-03-14T09:04:27.212014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1.preprocessing.RobustScaler()：使用具有鲁棒性的统计量缩放带有异常值（离群值）的数据，该缩放器删除中位数，并根据百分位数范围（默认值为IQR：四分位间距）缩放数据。 IQR是第1个四分位数（25%）和第3个四分位数（75%）之间的范围。数据集的标准是通过去除均值，缩放单位方差来完成，但是异常值通常会对样本的均值和方差造成负面影响，当异常值噪声很大时，用中位数和四分位数范围通常时候产生更坏的效果\n\n2.fit(): 用来计算mean（均值）和std（标准差），以便后面进行数据的标准化\ntransform(): 根据fit()函数计算的mean和std对数据进行标准化\nfit_transform(): 是fit()函数和transform()函数的组合，先进行fit，之后再进行transform（标准化）\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nFunctions for parsing and feature generation\n</div>","metadata":{}},{"cell_type":"code","source":"# The normalize function is required to normalize the data for the neural network.\n\ndef normalize(X_train, X_valid, X_test, normalize_opt, excluded_feat):\n    feats = [f for f in X_train.columns if f not in excluded_feat]\n    if normalize_opt != None:\n        if normalize_opt == 'min_max':\n            scaler = preprocessing.MinMaxScaler()#缩放到统一范围，【0.1】\n        elif normalize_opt == 'robust':\n            scaler = preprocessing.RobustScaler()\n        elif normalize_opt == 'standard':\n            scaler = preprocessing.StandardScaler()#标准化\n        elif normalize_opt == 'max_abs':\n            scaler = preprocessing.MaxAbsScaler()#归一化【-1,1】\n        scaler = scaler.fit(X_train[feats])\n        X_train[feats] = scaler.transform(X_train[feats])\n        X_valid[feats] = scaler.transform(X_valid[feats])\n        X_test[feats] = scaler.transform(X_test[feats])\n    return X_train, X_valid, X_test","metadata":{"execution":{"iopub.status.busy":"2023-03-17T09:57:50.279048Z","iopub.execute_input":"2023-03-17T09:57:50.279338Z","iopub.status.idle":"2023-03-17T09:57:50.28664Z","shell.execute_reply.started":"2023-03-17T09:57:50.279287Z","shell.execute_reply":"2023-03-17T09:57:50.284908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# functions for feature generation\n# Create random noise for robustness\nnp.random.seed(1337)\nnoise = np.random.normal(0, 0.5, 150_000)#其中的每个元素都是从均值为0、标准差为0.5的正态分布中随机采样得到的数值。\n#也就是说，这个数组包含了150,000个在[-0.5, 0.5]范围内的随机噪声值。\n\n# Mean Absolute Deviation  MAD（mean absolute deviation）又称为绝对值差中位数法\ndef maddest(d, axis=None):\n    return np.mean(np.absolute(d - np.mean(d, axis)), axis)#衡量一组数据离散程度的一种统计量\n\n# Denoise the raw signal given a segment x\ndef denoise_signal(x, wavelet='db4', level=1):\n    coeff = pywt.wavedec(x, wavelet, mode=\"per\")\n    sigma = (1/0.6745) * maddest(coeff[-level])\n    uthresh = sigma * np.sqrt(2*np.log(len(x)))\n    coeff[1:] = (pywt.threshold(i, value=uthresh, mode='hard') for i in coeff[1:])\n\n    return pywt.waverec(coeff, wavelet, mode='per')\n\n# Denoise the raw signal (simplified) given a segment x\ndef denoise_signal_simple(x, wavelet='db4', level=1):\n    coeff = pywt.wavedec(x, wavelet, mode=\"per\")\n    #univeral threshold\n    uthresh = 10\n    coeff[1:] = (pywt.threshold(i, value=uthresh, mode='hard') for i in coeff[1:])\n    # Reconstruct the signal using the thresholded coefficients\n    return pywt.waverec(coeff, wavelet, mode='per')\n\n# Generate the features given a segment z\ndef feature_gen(z):\n    X = pd.DataFrame(index=[0], dtype=np.float64)\n    \n    # Add noise, subtract median to remove bias from mean/median as time passes in the experiment\n    # Save the result as a new segment, z\n    z = z + noise#[11.64840635  5.75485882  7.83909284 ...  6.43139967  2.12669680.02331202]\n    #print('z:',z)\n    z = z - np.median(z)#[ 6.77510676  0.88155923  2.96579325 ...  1.55810008 -2.74660279  -4.84998757]\n    #print('z - np.median(z):',z)\n\n    # Save denoised versions of z\n    den_sample = denoise_signal(z)\n    #print('den_sample:',den_sample)#[0.1473848  0.14752262 0.14766243 ... 0.14701126 0.14711776 0.14724145]\n    den_sample_simple = denoise_signal_simple(z)\n    #print('den_sample_simple：',den_sample_simple)#[0.22098851 0.22114658 0.22130691 ... 0.22055999 0.22068216 0.22082405]\n    \n    # Mel-frequency cepstral coefficients\n    \n    mfcc = librosa.feature.mfcc(z)#mfcc函数可以用来提取音频的梅尔频率倒谱系数（Mel-Frequency Cepstral Coefficients，MFCCs）特征，MFCC被广泛应用于语音识别\n    #print('MFCC:',mfcc)\n    mfcc_mean = mfcc.mean(axis=1)\n    #print('MFCC_MEAN:',mfcc_mean)\n    mfcc_denoise_simple = librosa.feature.mfcc(den_sample_simple)\n    #print('mfcc_denoise_simple:',mfcc_denoise_simple)\n    mfcc_mean_denoise_simple = mfcc_denoise_simple.mean(axis=1) #0-19\n    \n    # Spectral contrast\n    lib_spectral_contrast_denoise_simple = librosa.feature.spectral_contrast(den_sample_simple).mean(axis=1) #0-6\n    lib_spectral_contrast = librosa.feature.spectral_contrast(z).mean(axis=1) #0-6\n    \n    # Neural network features\n    X['NN_zero_crossings_denoise'] = len(np.where(np.diff(np.sign(den_sample)))[0])#sign取数字前的正负号所以就是三个数{-1,0,1},#diff:a[n]-a[n-1]\n    X['NN_LGBM_percentile_roll20_std_50'] = np.percentile(pd.Series(z).rolling(20).std().dropna().values, 50)#滚动窗口尺寸50的标准差为20%，percentile是计算百分位数的\n    X['NN_q95_roll20_std'] = np.quantile(pd.Series(z).rolling(20).std().dropna().values, 0.95)\n    X['NN_LGBM_mfcc_mean4'] = mfcc_mean[4]#mfcc_mean[4] 表示 MFCC 系数矩阵中第 5 个系数在所有帧上的平均值，是一个实数。\n    X['NN_lib_spectral_contrast0'] = lib_spectral_contrast[0]\n    X['NN_num_peaks_3_denoise'] = feature_calculators.number_peaks(den_sample, 3)#计算时间序列中至少支持n的峰值数\n    X['NN_mfcc_mean_denoise_simple2'] = mfcc_mean_denoise_simple[2]\n    X['NN_mfcc_mean5'] = mfcc_mean[5]\n    X['NN_mfcc_mean2'] = mfcc_mean[2]\n    X['NN_mfcc_mean_denoise_simple5'] = mfcc_mean_denoise_simple[5]\n    X['NN_absquant95'] = np.quantile(np.abs(z), 0.95)#分位数0——1分位数\n    X['NN_median_roll50_std_denoise_simple'] = np.median(pd.Series(den_sample_simple).rolling(50).std().dropna().values)\n    X['NN_mfcc_mean_denoise_simple1'] = mfcc_mean_denoise_simple[1]\n    X['NN_quant99'] = np.quantile(z, 0.99)\n    X['NN_lib_zero_cross_rate_denoise_simple'] = librosa.feature.zero_crossing_rate(den_sample_simple)[0].mean()\n    X['NN_fftr_max_denoise'] = np.max(pd.Series(np.abs(np.fft.fft(den_sample)))[0:75000])\n    X['NN_abssumgreater15'] = np.sum(abs(z[np.where(abs(z)>15)]))\n    X['NN_LGBM_mfcc_mean18'] = mfcc_mean[18]\n    X['NN_lib_spectral_contrast_denoise_simple2'] = lib_spectral_contrast_denoise_simple[2]\n    X['NN_fftr_sum'] = np.sum(pd.Series(np.abs(np.fft.fft(z)))[0:75000])\n    X['NN_mfcc_mean_denoise_simple10'] = mfcc_mean_denoise_simple[10]\n    \n    # Extra features only LGBM used.\n    X['LGBM_num_peaks_2_denoise_simple'] = feature_calculators.number_peaks(den_sample_simple, 2)\n    X['LGBM_autocorr5'] = feature_calculators.autocorrelation(pd.Series(z), 5)#根据公式计算指定滞后的自相关\n    \n    # Windowed fast fourier transformations\n    fftrhann20000 = np.sum(np.abs(np.fft.fft(np.hanning(len(z))*z)[:20000]))\n    fftrhann20000_denoise = np.sum(np.abs(np.fft.fft(np.hanning(len(z))*den_sample)[:20000]))\n    fftrhann20000_diff_rate = (fftrhann20000 - fftrhann20000_denoise)/fftrhann20000\n    \n    X['LGBM_fftrhann20000_diff_rate'] = fftrhann20000_diff_rate\n    \n    return X","metadata":{"execution":{"iopub.status.busy":"2023-03-17T09:57:50.287868Z","iopub.execute_input":"2023-03-17T09:57:50.288321Z","iopub.status.idle":"2023-03-17T09:57:50.317347Z","shell.execute_reply.started":"2023-03-17T09:57:50.288268Z","shell.execute_reply":"2023-03-17T09:57:50.316668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nCode Explantion\n</div>","metadata":{}},{"cell_type":"markdown","source":"z:使得z的值在加入噪声、减去中位数等操作后更符合某些特定的要求或条件。这一步操作的目的是对z进行居中处理，即将z中每个数值向中位数靠拢，从而使得z的整体趋势更加明显，更符合某些特定的分析要求。最终得到的z是一个经过噪声、居中等处理的一维数组，可以用于某些特定的统计分析或建模任务。例如，可以用z来评估某种算法的鲁棒性，或者用z来训练一个能够抵抗噪声影响的机器学习模型。\n\ndef denoise_signal函数：具体来说，该算法接收一个一维的信号 x，并使用指定的小波基 wavelet 对其进行小波分解，分解级数为 level。然后，通过计算小波系数的中位绝对偏差（MAD，median absolute deviation），估算噪声的标准差 sigma，并将 sigma 与一个阈值 uthresh 进行比较，以决定哪些小波系数应该被阈值处理。最后，使用阈值处理后的小波系数重构信号，并返回去噪后的信号。\n具体的实现细节如下：\n使用 pywt.wavedec 函数对信号进行小波分解，得到一个包含多个小波系数的元组 coeff，其中 coeff[0] 是最低频的小波系数，而 coeff[i] 是第 i 层小波分解的小波系数。计算 coeff 中最后 level 层小波系数的 MAD，得到噪声的标准差 sigma。计算阈值 uthresh = sigma * np.sqrt(2*np.log(len(x)))，其中 len(x) 是信号的长度。对 coeff 中第 i 层小波系数进行硬阈值处理（hard thresholding），即将绝对值小于 uthresh 的小波系数设为 0，将绝对值大于等于 uthresh 的小波系数保留不变。使用 pywt.waverec 函数将阈值处理后的小波系数重构为去噪后的信号，返回去噪后的信号。需要注意的是，该算法使用了周期延拓模式（'per' mode）进行小波分解和重构，因此处理后的信号长度与原始信号长度相等\n\ndef denoise_signal_simple函数：该函数可以用于对输入信号进行简单的去噪处理。它使用小波变换将信号分解为多个尺度的系数，然后将这些系数阈值化以去除高频噪声。在这个简单版本中，阈值设置为一个常数10。最后，使用小波重构来生成去噪后的信号。\n\n\nNN_zero_crossings_denoise：因此，代码很可能是试图计算 den_sample 中分母符号变化的次数。具体来说，它首先计算 den_sample 中每个元素的符号，然后取相邻符号之间的差异，最后计算符号变化的次数（即 np.diff(np.sign(den_sample)) 非零的次数）。结果表示某个分数的分母在 den_sample 中变号的次数，这可能与某些情况有关（例如，分析函数或数据集的行为）。\n\nNN_LGBM_percentile_roll20_std_50：因此，整个代码的含义是：对变量 z 中的数据创建一个 Pandas 的 Series 对象，然后计算其滚动窗口内每个子序列的标准差，并删除含有缺失值的行。接着将其作为一个 NumPy 数组，并计算其在第 50 个百分位数处的值。换句话说，这段代码的作用是计算 z 的滚动窗口内标准差的中位数。\n\nNN_q95_roll20_std：因此，整个代码的含义是：对变量 z 中的数据创建一个 Pandas 的 Series 对象，然后计算其滚动窗口内每个子序列的标准差，并删除含有缺失值的行。接着将其作为一个 NumPy 数组，并计算其在第 95 个百分位数处的值。换句话说，这段代码的作用是计算 z 的滚动窗口内标准差的第 95 个百分位数（即 0.95 分位数）。\n\n具体来说，librosa.feature.mfcc(z) 首先将输入的音频信号进行梅尔频率变换，然后将结果用离散余弦变换（DCT）转换为一组 MFCC 系数。输出结果是一个二维 NumPy 数组，其中每一行代表一个 MFCC 系数，每一列代表输入音频信号的一个帧。\n\nMFCC 全称是梅尔频率倒谱系数（Mel-Frequency Cepstral Coefficients），是一种音频信号的特征表示方法，主要用于音频处理领域，例如音频信号分类、语音识别、音乐信息检索等任务。\n\nMFCC 最初是为了模拟人耳对声音的感知而设计的。我们的耳朵并不是等响度的，而是能够感知到较高频率的信号要比较低频率的信号更加敏感。为了模拟这种非线性的频率响应，MFCC 使用了梅尔滤波器组来模拟人耳的频率响应特征。梅尔滤波器组可以将音频信号在频域上进行滤波，使得低频段的频率细节被更加精细地刻画，而高频段的频率细节则被更加粗略地刻画。\nMFCC 主要分为以下几个步骤：\n预加重：对输入的音频信号进行预加重处理，可以通过滤波器对音频信号进行高通滤波，去除信号中的直流分量和低频部分，使得信号更加平滑。\n帧分割：将音频信号分割成多个时间窗口，每个时间窗口称为一帧，通常采用重叠窗口来提高帧之间的连续性。\n傅里叶变换：对每一帧的音频信号进行离散傅里叶变换（DFT），将信号从时域转换到频域。\n梅尔滤波器组：将每一帧的信号通过一组梅尔滤波器，将信号在频域上进行滤波。\n对数变换：将每一帧的信号在滤波器组之后，进行对数变换，即取每个滤波器的对数值。这个过程可以模拟人耳对声音的感知，使得高频部分被更好地表达。\n离散余弦变换：对每一帧的信号进行离散余弦变换（DCT），得到一组 MFCC 系数。MFCC 系数可以认为是对信号频谱信息的一种压缩和表示。\n能量归一化：对每一帧的 MFCC 系数进行能量归一化，使得不同帧之间的能量尺度保持一致。\nMFCC 能够很好地表达音频信号的频谱信息，对于许多音频处理任务都有很好的应用效果。\n\nlibrosa.feature.mfcc(z) 表示使用 Librosa 库中的 mfcc 函数计算音频信号 z 的 MFCC 特征表示。这个过程包括了预处理、帧分割、傅里叶变换、梅尔滤波器组、对数变换、离散余弦变换等步骤，最终输出一个 MFCC 系数矩阵。\nmfcc.mean(axis=1) 表示对 MFCC 系数矩阵中每一行（即每个 MFCC 系数）沿着第二个维度进行均值计算，得到一个长度为 n_mfcc 的向量 mfcc_mean，其中 n_mfcc 是 MFCC 系数的数量。这个向量可以作为音频信号的一个紧凑的特征表示，可以用于后续的音频处理任务，例如音频分类、语音识别等。\n\nfeature_calculators.number_peaks 表示使用 tsfresh 库中的 number_peaks 函数计算时间序列中峰值的数量。den_sample 是一个一维的时间序列，可以是一段音频信号的强度波形、某个传感器采集到的物理量、某个时间序列数据等。3 是一个阈值，表示只有高度大于等于 3 的峰值才会被计算在峰值数量中。具体来说，对于任意一个峰值，如果它的高度小于 3，则不会被计入峰值数量中；如果它的高度大于等于 3，则会被计入峰值数量中。\n\nNN_fftr_max_denoise：np.fft.fft(den_sample) 表示对时间序列 den_sample 进行快速傅里叶变换，得到其频域表示。傅里叶变换可以将一个时间域的信号转换为一个频域的信号，表示信号在不同频率上的能量分布。np.abs() 表示计算傅里叶变换结果的振幅谱，即傅里叶变换结果的模值。pd.Series() 表示将振幅谱转换为 Pandas 库中的 Series 对象，方便进行后续的数据处理。[0:75000] 表示选取前 75000 个频率分量的振幅。np.max() 表示对前 75000 个频率分量的振幅序列取最大值，即得到一个标量，表示该时间序列的振幅谱中最大的前 75000 个频率分量的振幅。\n\nNN_abssumgreater15：np.where(abs(z)>15) 表示返回一个布尔型数组，其元素与 abs(z) 中的元素一一对应，元素值为 True 表示该位置上的元素绝对值大于 15，否则为 False。z[np.where(abs(z)>15)] 表示根据上一步的布尔型数组，从数组 z 中选取所有满足条件的元素，组成一个新的数组abs(z[np.where(abs(z)>15)]) 表示将新数组中的所有元素取绝对值。np.sum(abs(z[np.where(abs(z)>15)])) 表示对新数组中的所有元素取绝对值之后，求它们的和，即得到一个标量，表示数组 z 中绝对值大于 15 的元素的绝对值之和。\n\nfeature_calculators.autocorrelation(pd.Series(z), 5) 表示调用函数计算数组 z 的自相关函数在时延 5 处的取值，即得到一个标量，表示数组 z 在时延为 5 时的自相关函数值。\n\n\nnp.hanning(len(z)) 表示用汉宁窗对长度为 len(z) 的一维数组 z 进行加窗处理。在信号处理中，为了避免傅里叶变换（或其他频域分析方法）时的频谱泄露问题，通常需要对信号进行加窗处理。加窗可以使信号在边缘处衰减，从而减小边缘处的波动，降低噪声对频域分析结果的影响。\nnp.hanning(len(z))*z 表示用汉宁窗对数组 z 进行加窗处理。\nnp.fft.fft(np.hanning(len(z))*z) 表示对加窗后的数组 z 进行傅里叶变换，并返回变换后的复数数组。\nnp.abs(np.fft.fft(np.hanning(len(z))*z)) 表示对傅里叶变换后的复数数组取绝对值，得到幅度谱（amplitude spectrum）。\nnp.sum(np.abs(np.fft.fft(np.hanning(len(z))*z)[:20000])) 表示计算幅度谱的前 20000 个系数的总和。\nfftrhann20000_denoise 表示对 den_sample 应用与步骤 1-5 相同的计算过程，计算出其前 20000 个系数的幅度谱和总和。\nfftrhann20000_diff_rate 表示用 fftrhann20000 减去 fftrhann20000_denoise 得到一个差值，再除以 fftrhann20000，计算出两个幅度谱之间的比率变化率。这个值可以用来衡量 den_sample 与 z 的频域信息之间的差异程度。\n\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nCode Output\n</div>","metadata":{}},{"cell_type":"markdown","source":"**********这是第***次调用**********\nz: [11.64840635  5.75485882  7.83909284 ...  6.43139967  2.1266968\n  0.02331202]\nz - np.median(z): [ 6.77510676  0.88155923  2.96579325 ...  1.55810008 -2.74660279\n -4.84998757]\nden_sample: [0.1473848  0.14752262 0.14766243 ... 0.14701126 0.14711776 0.14724145]\nden_sample_simple： [0.22098851 0.22114658 0.22130691 ... 0.22055999 0.22068216 0.22082405]\nMFCC: [[ 3.17551290e+02  3.40446656e+02  3.73843809e+02 ...  3.51661158e+02\n   3.41824124e+02  3.20568114e+02]\n [ 9.14700992e+00  2.85420534e+01  5.90111755e+01 ...  3.82318881e+01\n   3.12573068e+01  1.47065117e+01]\n [-3.03439805e+01 -3.38080384e+01 -2.74848621e+01 ... -3.93928879e+01\n  -3.21175624e+01 -2.41329995e+01]\n ...\n [-1.09451379e+00 -1.32636709e+00 -3.73579774e-01 ...  1.06833677e+00\n   3.46787427e+00  2.29846504e-01]\n [-7.35424316e+00 -1.91377845e+00  2.04084144e+00 ... -1.25641415e+00\n  -5.71275400e-01 -3.24189787e+00]\n [-9.89417955e-02 -2.52687166e-01 -1.15438393e+00 ...  8.17363750e-01\n   1.53031132e+00 -9.54438483e-01]]\nMFCC_MEAN: [ 3.35021974e+02  2.66549004e+01 -2.79846261e+01 -1.82822281e+01\n -1.13690578e+01  5.27948579e+00 -2.82513608e-01 -2.65241459e+00\n -4.66760263e+00  7.78315810e-01  2.47239883e+00  1.43491510e+00\n -3.84506463e+00 -3.22263540e+00  2.35090147e+00  6.74574048e+00\n  2.93653081e+00 -1.92530747e+00 -1.77808754e+00  1.51449941e+00]\nmfcc_denoise_simple: [[ 1.47908131e+02  2.31719427e+02  3.04166241e+02 ...  2.68672230e+02\n   2.23478440e+02  1.28425439e+02]\n [ 8.61641503e+01  1.02932093e+02  1.25405630e+02 ...  8.57150734e+01\n   9.22431146e+01  9.37996143e+01]\n [-1.24434767e+02 -8.55143526e+01 -7.18673461e+01 ... -9.85708759e+01\n  -1.00445882e+02 -1.17257751e+02]\n ...\n [ 8.26055586e+00  1.48097249e+00 -9.95883655e-01 ... -8.10694307e-01\n   1.48046218e+00  5.48782634e+00]\n [-4.04372613e+00 -6.93758306e+00 -4.94788554e+00 ... -7.87799562e+00\n  -4.08349441e+00  3.19111994e+00]\n [-6.96935427e-01 -2.24644537e-01  4.51786719e+00 ...  3.86441104e+00\n   4.81516406e+00  6.35623690e+00]]\n**********这是第***次调用**********\nz: [4.64840635 5.75485882 7.83909284 ... 7.43139967 7.1266968  5.02331202]\nz - np.median(z): [-0.07301813  1.03343434  3.11766836 ...  2.70997519  2.40527232\n  0.30188754]\nden_sample: [-0.06020885 -0.06047565 -0.06074633 ... -0.05948654 -0.05969196\n -0.05993111]\nden_sample_simple： [-0.01001956 -0.0101286  -0.01023926 ... -0.00972402 -0.00980837\n -0.00990623]\nMFCC: [[ 3.03863300e+02  3.13279475e+02  3.13374210e+02 ...  3.27389832e+02\n   3.20460787e+02  3.13226703e+02]\n [ 3.26613861e+00  9.61141154e+00  8.90234592e+00 ...  2.35959280e+01\n   1.89973727e+01  1.79307311e+01]\n [-2.70761642e+01 -2.40863431e+01 -2.02333033e+01 ... -2.74425399e+01\n  -2.41400667e+01 -1.83804899e+01]\n ...\n [ 2.21037871e+00 -1.04406161e+00 -4.76308578e+00 ... -2.54038379e+00\n  -7.96586718e-01  1.16624384e-01]\n [-1.17887292e+00 -5.27833965e+00 -6.75139729e+00 ...  1.09861866e+00\n  -1.21394819e+00 -2.26537869e-01]\n [-2.01919969e+00 -2.11477714e-01 -5.65513944e-01 ...  1.18509859e+00\n   5.34879056e+00  5.13511006e+00]]\nMFCC_MEAN: [ 3.37351418e+02  2.81633652e+01 -2.84358372e+01 -1.85306690e+01\n -1.10275357e+01  5.63720496e+00 -5.05431258e-01 -2.98677999e+00\n -5.11821226e+00  2.83126942e-01  2.98090413e+00  1.98164583e+00\n -3.71423526e+00 -3.69425847e+00  2.26685356e+00  6.62478627e+00\n  3.02082327e+00 -2.10346361e+00 -2.46446200e+00  5.16791351e-01]\nmfcc_denoise_simple: [[ 161.66243416  140.42327502   80.57946109 ...  172.74095973\n   133.43738075   54.20272351]\n [  10.79338534   27.72984067   68.53720477 ...  102.98256214\n    82.18837734   43.4856973 ]\n [-101.54735639 -101.51914868 -105.61703768 ...  -98.52567999\n  -107.96827722 -111.98945084]\n ...\n [   3.9411515     4.79417074    5.96463    ...    1.16224626\n     1.96893998    1.04489455]\n [   4.98784465    6.45042773    8.64851177 ...    3.98041136\n     6.89937777    6.80065343]\n [   0.90778179    3.25227325    0.9802916  ...   -0.97916514\n     0.80571427   -2.7585528 ]]\n**********这是第***次调用**********\nz: [4.64840635 4.75485882 7.83909284 ... 5.43139967 8.1266968  3.02331202]\nz - np.median(z): [-0.24207566 -0.13562319  2.94861083 ...  0.54091766  3.2362148\n -1.86716999]\nden_sample: [0.11733676 0.11938673 0.12572997 ... 0.14676823 0.12184574 0.1132612 ]\nden_sample_simple： [0.58450351 0.67334034 0.77290528 ... 0.39672531 0.43461733 0.49821934]\nMFCC: [[310.38704996 311.57683107 325.56373847 ... 319.70194811 320.38700305\n  326.59376639]\n [ 10.54021543   9.79290854  15.50871981 ...  19.93019327  14.37070858\n   18.15904765]\n [-17.10550435 -17.65878455 -24.29474325 ... -17.93331995 -19.32696825\n  -24.51294215]\n ...\n [  0.89727033   4.2432036    0.49888222 ...  -0.99294478   1.78242825\n   -2.13543316]\n [  2.13347381   1.59672534  -3.30802185 ...  -2.05353102  -2.09887767\n   -4.1850144 ]\n [  0.82292303   2.2706763    0.77016951 ...   5.28400155   2.63543216\n    1.28095936]]\nMFCC_MEAN: [ 3.43276577e+02  3.34197987e+01 -2.83212161e+01 -2.13807132e+01\n -1.28739701e+01  5.27431739e+00 -1.19098766e+00 -3.90433815e+00\n -5.72509815e+00  2.36090609e-01  3.15897578e+00  1.96499011e+00\n -4.40864399e+00 -4.56637673e+00  2.08708567e+00  7.45077056e+00\n  3.11828461e+00 -2.40438594e+00 -2.72190178e+00  8.21391731e-01]\nmfcc_denoise_simple: [[  76.31019072   56.2047817   133.43285426 ...  119.10720896\n    98.28665755  124.067528  ]\n [  90.55558664   65.53389443   80.52528722 ...   97.44127072\n    83.74569782   88.61284779]\n [-104.83051642 -117.78771647 -128.17205163 ... -115.82642726\n  -119.99600039 -120.64400659]\n ...\n [   6.8760496     5.50359959    5.11979229 ...    2.0988799\n     7.47656614    5.60936578]\n [   8.02944578    9.57887141   -8.02532688 ...    7.28919513\n     5.78160255    3.71539697]\n [   1.0948506    -0.97890074    3.22924746 ...    9.37112242\n     5.7415966     6.54622963]]","metadata":{}},{"cell_type":"code","source":"# create train and test sets\n\ndef parse_sample(sample, start):\n\n    #print(f'**********这是第***次调用**********')\n    delta = feature_gen(sample['acoustic_data'].values)\n    delta['start'] = start\n    delta['target'] = sample['time_to_failure'].values[-1]\n\n    return delta\n    \ndef sample_train_gen(df, segment_size=150_000, indices_to_calculate=[0]):\n    result = Parallel(n_jobs=1, temp_folder=\"/tmp\", max_nbytes=None, backend=\"multiprocessing\")(delayed(parse_sample)(df[int(i) : int(i) + segment_size], int(i)) \n                                                                                                for i in tqdm(indices_to_calculate))\n    data = [r.values for r in result]\n    data = np.vstack(data)\n    X = pd.DataFrame(data, columns=result[0].columns)\n    X = X.sort_values(\"start\")\n    return X\n\ndef parse_sample_test(seg_id):\n    sample = pd.read_csv('../input/test/' + seg_id + '.csv', dtype={'acoustic_data': np.int32})\n    delta = feature_gen(sample['acoustic_data'].values)\n    delta['seg_id'] = seg_id\n    return delta\n\ndef sample_test_gen():\n    X = pd.DataFrame()\n    submission = pd.read_csv('../input/sample_submission.csv', index_col='seg_id')\n    result = Parallel(n_jobs=1, temp_folder=\"/tmp\", max_nbytes=None, backend=\"multiprocessing\")(delayed(parse_sample_test)(seg_id) for seg_id in tqdm(submission.index))\n    data = [r.values for r in result]\n    data = np.vstack(data)\n    X = pd.DataFrame(data, columns=result[0].columns)\n    return X\n\nindices_to_calculate = raw.index.values[::150_000][:-1]\n#print('indices_to_calculate:',indices_to_calculate)\ntrain = sample_train_gen(raw, indices_to_calculate=indices_to_calculate)\ndel raw\ngc.collect()\ntest = sample_test_gen()","metadata":{"execution":{"iopub.status.busy":"2023-03-17T09:57:50.318945Z","iopub.execute_input":"2023-03-17T09:57:50.319495Z","iopub.status.idle":"2023-03-17T10:28:37.518021Z","shell.execute_reply.started":"2023-03-17T09:57:50.319369Z","shell.execute_reply":"2023-03-17T10:28:37.517183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test","metadata":{"execution":{"iopub.status.busy":"2023-03-16T10:03:41.130987Z","iopub.execute_input":"2023-03-16T10:03:41.131273Z","iopub.status.idle":"2023-03-16T10:03:41.224427Z","shell.execute_reply.started":"2023-03-16T10:03:41.131219Z","shell.execute_reply":"2023-03-16T10:03:41.223487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2023-03-14T09:35:47.18937Z","iopub.execute_input":"2023-03-14T09:35:47.189684Z","iopub.status.idle":"2023-03-14T09:35:47.283952Z","shell.execute_reply.started":"2023-03-14T09:35:47.189631Z","shell.execute_reply":"2023-03-14T09:35:47.283247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare Cross-Validation","metadata":{}},{"cell_type":"code","source":"# keep CV observations from train set\netq_meta = [\n{\"start\":0,         \"end\":5656574},\n{\"start\":5656574,   \"end\":50085878},\n{\"start\":50085878,  \"end\":104677356},\n{\"start\":104677356, \"end\":138772453},\n{\"start\":138772453, \"end\":187641820},\n{\"start\":187641820, \"end\":218652630},\n{\"start\":218652630, \"end\":245829585},\n{\"start\":245829585, \"end\":307838917},\n{\"start\":307838917, \"end\":338276287},\n{\"start\":338276287, \"end\":375377848},\n{\"start\":375377848, \"end\":419368880},\n{\"start\":419368880, \"end\":461811623},\n{\"start\":461811623, \"end\":495800225},\n{\"start\":495800225, \"end\":528777115},\n{\"start\":528777115, \"end\":585568144},\n{\"start\":585568144, \"end\":621985673},\n{\"start\":621985673, \"end\":629145480},\n]\n\nfor i, etq in enumerate(etq_meta):\n    train.loc[(train['start'] + 150_000 >= etq[\"start\"]) & (train['start'] <= etq[\"end\"] - 150_000), \"eq\"] = i\n\n# We are only keeping segments that belong in these earthquakes\n# This is to make the training distribution more like the testing distribution\n#如果上述两个条件都为 True，则将 train 中相应数据点的 eq 列的值设置为 i，即当前地震事件的索引。\n#这样，train 中每个数据点的 eq 列都会被设置为其所属的地震事件的索引。\ntrain_sample = train[train[\"eq\"].isin([2, 7, 0, 4, 11, 13, 9, 1, 14, 10])]#17行","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:37.519474Z","iopub.execute_input":"2023-03-17T10:28:37.519817Z","iopub.status.idle":"2023-03-17T10:28:37.665997Z","shell.execute_reply.started":"2023-03-17T10:28:37.519742Z","shell.execute_reply":"2023-03-17T10:28:37.665122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample","metadata":{"execution":{"iopub.status.busy":"2023-03-16T09:55:48.699116Z","iopub.execute_input":"2023-03-16T09:55:48.699406Z","iopub.status.idle":"2023-03-16T09:55:48.794975Z","shell.execute_reply.started":"2023-03-16T09:55:48.699352Z","shell.execute_reply":"2023-03-16T09:55:48.794214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# delete unnecessary files\ndel train\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:37.667146Z","iopub.execute_input":"2023-03-17T10:28:37.667404Z","iopub.status.idle":"2023-03-17T10:28:37.76722Z","shell.execute_reply.started":"2023-03-17T10:28:37.667359Z","shell.execute_reply":"2023-03-17T10:28:37.766113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# reset the index of the final train set\ntrain_sample=train_sample.reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:37.769133Z","iopub.execute_input":"2023-03-17T10:28:37.769494Z","iopub.status.idle":"2023-03-17T10:28:37.774904Z","shell.execute_reply.started":"2023-03-17T10:28:37.769415Z","shell.execute_reply":"2023-03-17T10:28:37.773813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample","metadata":{"execution":{"iopub.status.busy":"2023-03-16T09:56:43.572234Z","iopub.execute_input":"2023-03-16T09:56:43.572581Z","iopub.status.idle":"2023-03-16T09:56:43.673126Z","shell.execute_reply.started":"2023-03-16T09:56:43.572518Z","shell.execute_reply":"2023-03-16T09:56:43.672417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create time since failure target variable\n# This will be used in the NN as an additional objective\ntargets=train_sample[['target','start']]\ntargets['tsf']=targets['target']-targets['target'].shift(1).fillna(0)\ntargets['tsf']=np.where(targets['tsf']>1.5, targets['tsf'], 0)\ntargets['tsf'].iloc[0]=targets['target'].iloc[0]\n\ntemp_max=0\nfor i in tqdm(range(targets.shape[0])):\n    if targets['tsf'].iloc[i]>0:\n        temp_max=targets['tsf'].iloc[i]\n    else:\n        targets['tsf'].iloc[i]=temp_max\n        \ntargets['tsf']=targets['tsf']-targets['target']\n\n# create a flag target variable for TTF<0.5 secs\n# This will be used in the NN as an additional objective\ntarget=targets['target'].copy().values\ntarget[target>=0.5]=1\ntarget[target<0.5]=0\ntarget=1-target\n\ntargets['binary']=target\ndel target\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:37.776391Z","iopub.execute_input":"2023-03-17T10:28:37.777098Z","iopub.status.idle":"2023-03-17T10:28:38.426537Z","shell.execute_reply.started":"2023-03-17T10:28:37.776836Z","shell.execute_reply":"2023-03-17T10:28:38.425761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"targets","metadata":{"execution":{"iopub.status.busy":"2023-03-16T09:58:55.271126Z","iopub.execute_input":"2023-03-16T09:58:55.271423Z","iopub.status.idle":"2023-03-16T09:58:55.303165Z","shell.execute_reply.started":"2023-03-16T09:58:55.271368Z","shell.execute_reply":"2023-03-16T09:58:55.302412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"这段代码的作用是创建一个名为tsf的新列，用于存储target列数据的变化量。具体来说，这行代码使用了shift函数将target列的数据向上平移一行，然后用平移后的结果减去原始的target列数据，得到相邻两行数据之间的差值。这样可以计算出target列中每一行数据与前一行数据之间的差异，也就是它们的变化量。由于在数据的第一行没有前一行数据，因此shift函数返回的第一行数据会是NaN。为了确保代码的正确性，这行代码使用了fillna(0)函数，将NaN值替换为0。这样可以保证在计算变化量时，第一行数据的变化量为当前数据值减去0，即当前数据值本身。\n\n即相邻两行数据之间的差值。如果两行数据之间的变化量大于1.5，就将该变化量保留在tsf列中，否则将tsf列设为0。","metadata":{}},{"cell_type":"code","source":"# import submission file\nsubmission = pd.read_csv('../input/sample_submission.csv', index_col='seg_id')","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:38.427809Z","iopub.execute_input":"2023-03-17T10:28:38.428285Z","iopub.status.idle":"2023-03-17T10:28:38.438805Z","shell.execute_reply.started":"2023-03-17T10:28:38.42823Z","shell.execute_reply":"2023-03-17T10:28:38.437865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# delete unnecessary columns\ntrain_sample.drop(['start', 'target', 'eq'],axis=1,inplace=True)\ntest.drop(['seg_id'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:38.440252Z","iopub.execute_input":"2023-03-17T10:28:38.440688Z","iopub.status.idle":"2023-03-17T10:28:38.447947Z","shell.execute_reply.started":"2023-03-17T10:28:38.440481Z","shell.execute_reply":"2023-03-17T10:28:38.446976Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We also need to convert test columns from objects to float64\ntest = test.astype('float64')","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:38.449235Z","iopub.execute_input":"2023-03-17T10:28:38.449795Z","iopub.status.idle":"2023-03-17T10:28:38.461095Z","shell.execute_reply.started":"2023-03-17T10:28:38.449742Z","shell.execute_reply":"2023-03-17T10:28:38.460347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample","metadata":{"execution":{"iopub.status.busy":"2023-03-16T10:41:52.908798Z","iopub.execute_input":"2023-03-16T10:41:52.909089Z","iopub.status.idle":"2023-03-16T10:41:52.999628Z","shell.execute_reply.started":"2023-03-16T10:41:52.909037Z","shell.execute_reply":"2023-03-16T10:41:52.99872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test","metadata":{"execution":{"iopub.status.busy":"2023-03-16T10:42:00.095996Z","iopub.execute_input":"2023-03-16T10:42:00.096304Z","iopub.status.idle":"2023-03-16T10:42:00.183545Z","shell.execute_reply.started":"2023-03-16T10:42:00.096247Z","shell.execute_reply":"2023-03-16T10:42:00.182564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define your kfold cross validation\n# We used 3 folds, because we did not see improvements with higher folds\n# We are not scared of shuffling because The whole point of this comp is to be independent of time. Test is shuffled\nn_fold = 3\n\nkf = KFold(n_splits=n_fold, shuffle=True, random_state=1337)\nkf = list(kf.split(np.arange(len(train_sample))))","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:38.463869Z","iopub.execute_input":"2023-03-17T10:28:38.464163Z","iopub.status.idle":"2023-03-17T10:28:38.472166Z","shell.execute_reply.started":"2023-03-17T10:28:38.464102Z","shell.execute_reply":"2023-03-17T10:28:38.471289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nTrain the LGBM\n</div>\nThe LGBM is trained using only six features. We train using shuffled 3Fold. The LGBM is averaged over ten runs to improve generalization. Hyperparameters were optimized for cross-validation.","metadata":{}},{"cell_type":"code","source":"LGBM_feats = [feat for feat in train_sample.columns if 'LGBM' in feat]\nprint('The features LGBM is using are:', LGBM_feats)","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:38.473438Z","iopub.execute_input":"2023-03-17T10:28:38.473956Z","iopub.status.idle":"2023-03-17T10:28:38.481261Z","shell.execute_reply.started":"2023-03-17T10:28:38.47371Z","shell.execute_reply":"2023-03-17T10:28:38.480309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oof_LGBM = np.zeros(len(train_sample))\nsub_LGBM = np.zeros(len(submission))\nseeds = [0,1,2,3,4,5,6,7,8,9]\n\nfor seed in seeds:\n    print('Seed',seed)\n    for fold_n, (train_index, valid_index) in enumerate(kf):\n        print('Fold', fold_n)\n\n        # Create train and validation data using only LGBM_feats.\n        trn_data = lgb.Dataset(train_sample[LGBM_feats].iloc[train_index], label=targets['target'].iloc[train_index])\n        val_data = lgb.Dataset(train_sample[LGBM_feats].iloc[valid_index], label=targets['target'].iloc[valid_index])\n\n        params = {'num_leaves': 4, # Low number of leaves reduces LGBM complexity\n          'min_data_in_leaf': 5,\n          'objective':'fair', # Fitting to fair objective performed better than fitting to MAE objective\n          'max_depth': -1,\n          'learning_rate': 0.01,\n          \"boosting\": \"gbdt\", \n          'boost_from_average': True,\n          \"feature_fraction\": 0.9,\n          \"bagging_freq\": 1,\n          \"bagging_fraction\": 0.5,\n          \"bagging_seed\": 0,\n          \"metric\": 'mae',\n          \"verbosity\": -1,\n          'max_bin': 500,\n          'reg_alpha': 0, \n          'reg_lambda': 0,\n          'seed': seed,\n          'n_jobs': 1\n          }\n\n        clf = lgb.train(params, trn_data, 1000000, valid_sets = [trn_data, val_data], verbose_eval=1000, early_stopping_rounds = 1000)\n\n        oof_LGBM[valid_index] += clf.predict(train_sample[LGBM_feats].iloc[valid_index], num_iteration=clf.best_iteration)\n        sub_LGBM += clf.predict(test[LGBM_feats], num_iteration=clf.best_iteration) / n_fold\n        \noof_LGBM = oof_LGBM / len(seeds)\nsub_LGBM = sub_LGBM / len(seeds)\n    \nprint('\\nMAE for LGBM: ', mean_absolute_error(targets['target'], oof_LGBM))","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:28:38.482573Z","iopub.execute_input":"2023-03-17T10:28:38.483076Z","iopub.status.idle":"2023-03-17T10:29:07.825178Z","shell.execute_reply.started":"2023-03-17T10:28:38.483025Z","shell.execute_reply":"2023-03-17T10:29:07.824328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nTrain the NN\n</div>\n\nThe NN is trained using shuffled 3Fold. It is averaged over 8 runs to improve generalization. Sometimes when running the model, the initial weights are bad which results in bad results in cross-validation. If this happens, we will not use it when we average.\n\nThe model is simultaneously fit to three targets: Time to Failure (TTF), Time Since Failure (TSF), and Binary Target for TTF < 0.5 seconds. The loss weights are 8, 1, and 1, respectively. Because the NN has to focus on the TSF and Binary targets, the weights created seem to be better for predicting TTF. Likely, by fitting the NN this way, it reduces overfitting and increases generalization.\n\nWe use Nadam optimizer.","metadata":{}},{"cell_type":"code","source":"NN_feats = [feat for feat in train_sample.columns if 'NN' in feat]\nprint('The features NN is using are:', NN_feats)","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:29:07.826473Z","iopub.execute_input":"2023-03-17T10:29:07.826764Z","iopub.status.idle":"2023-03-17T10:29:07.832368Z","shell.execute_reply.started":"2023-03-17T10:29:07.826716Z","shell.execute_reply":"2023-03-17T10:29:07.831577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Subset columns to only use the neural network features\ntrain_sample = train_sample[NN_feats]\ntest = test[NN_feats]","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:29:07.833748Z","iopub.execute_input":"2023-03-17T10:29:07.834359Z","iopub.status.idle":"2023-03-17T10:29:07.846255Z","shell.execute_reply.started":"2023-03-17T10:29:07.834304Z","shell.execute_reply":"2023-03-17T10:29:07.845378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Define Neural Network architecture\ndef get_model():\n\n    inp = Input(shape=(1,train_sample.shape[1]))\n    x = BatchNormalization()(inp)\n    #x = LSTM(128,return_sequences=True)(x) # LSTM as first layer performed better than Dense.\n    x = GRU(128,return_sequences=True)(x)\n    x = Convolution1D(128, (2),activation='relu', padding=\"same\")(x)\n    x = Convolution1D(84, (2),activation='relu', padding=\"same\")(x)\n    x = Convolution1D(64, (2),activation='relu', padding=\"same\")(x)\n\n    x = Flatten()(x)\n\n    x = Dense(64, activation=\"relu\")(x)\n    x = Dense(32, activation=\"relu\")(x)\n    \n    #outputs\n    ttf = Dense(1, activation='relu',name='regressor')(x) # Time to Failure\n    tsf = Dense(1)(x) # Time Since Failure\n    classifier = Dense(1, activation='sigmoid')(x) # Binary for TTF<0.5 seconds\n    \n    model = models.Model(inputs=inp, outputs=[ttf,tsf,classifier])    \n    opt = optimizers.Nadam(lr=0.008)\n\n    # We are fitting to 3 targets simultaneously: Time to Failure (TTF), Time Since Failure (TSF), and Binary for TTF<0.5 seconds\n    # We weight the model to optimize heavily for TTF\n    # Optimizing for TSF and Binary TTF<0.5 helps to reduce overfitting, and helps for generalization.\n    model.compile(optimizer=opt, loss=['mae','mae','binary_crossentropy'],loss_weights=[8,1,1],metrics=['mae'])\n    return model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_model_2():\n    \n    inp = Input(shape=(1,train_sample.shape[1]))\n    x = BatchNormalization()(inp)\n    x = LSTM(128,return_sequences=True)(x)\n\n    x_conv = Convolution1D(128, (2),activation='relu', padding=\"same\")(x)\n    x_conv = GaussianDropout(0.1)(x_conv)\n    x_conv = Convolution1D(256, (2),activation='relu', padding=\"same\")(x_conv)\n    x_conv = Flatten()(x_conv)\n    \n    x = Dense(256, activation=\"relu\")(inp)\n    x = GaussianDropout(0.1)(x)\n    x = Reshape([16,16,1])(x)\n    x_ = Convolution2D(4, (5,5), activation=\"relu\", dilation_rate=3, padding=\"same\")(x)\n    x = Convolution2D(4, (7,7), activation=\"relu\", padding=\"same\")(x)\n    x = Concatenate()([x_, x])\n    x = GaussianDropout(0.1)(x)\n    x = Convolution2D(4, (3,3), activation=\"relu\", padding=\"same\")(x)\n    x = Flatten()(x)\n    \n    x = Concatenate()([x, x_conv])\n    \n    x = Dense(128, activation=\"relu\")(x)\n#     x = BatchNormalization()(x)\n     #outputs\n    ttf = Dense(1, activation='relu',name='regressor')(x) # Time to Failure\n    tsf = Dense(1)(x) # Time Since Failure\n    classifier = Dense(1, activation='sigmoid')(x) # Binary for TTF<0.5 seconds\n    \n    model = models.Model(inputs=inp, outputs=[ttf,tsf,classifier])    \n    opt = optimizers.Nadam(lr=0.008)\n\n    # We are fitting to 3 targets simultaneously: Time to Failure (TTF), Time Since Failure (TSF), and Binary for TTF<0.5 seconds\n    # We weight the model to optimize heavily for TTF\n    # Optimizing for TSF and Binary TTF<0.5 helps to reduce overfitting, and helps for generalization.\n    model.compile(optimizer=opt, loss=['mae','mae','binary_crossentropy'],loss_weights=[8,1,1],metrics=['mae'])\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:33:41.412664Z","iopub.execute_input":"2023-03-17T10:33:41.412963Z","iopub.status.idle":"2023-03-17T10:33:41.42088Z","shell.execute_reply.started":"2023-03-17T10:33:41.412908Z","shell.execute_reply":"2023-03-17T10:33:41.420094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n=8 # number of NN runs\n\noof_final = np.zeros(len(train_sample))\nsub_final = np.zeros(len(submission))\ni=0\n\n\nwhile i<8:\n    print('Running Model ', i+1)\n    \n    oof = np.zeros(len(train_sample))\n    prediction = np.zeros(len(submission))\n\n    for fold_n, (train_index, valid_index) in enumerate(kf):\n        #define training and validation sets\n\n        train_x=train_sample.iloc[train_index] #training set\n        train_y_ttf=targets['target'].iloc[train_index] #training target(Time to Failure)\n\n        valid_x=train_sample.iloc[valid_index] #validation set\n        valid_y_ttf=targets['target'].iloc[valid_index] #validation target(Time to Failure)\n\n        train_y_tsf=targets['tsf'].iloc[train_index] #training target(Time Since Failure)\n        train_y_clf=targets['binary'].iloc[train_index] #training target(Binary for TTF<0.5 Secs)\n\n        valid_y_tsf=targets['tsf'].iloc[valid_index] #validation target(Time Since Failure)\n        valid_y_clf=targets['binary'].iloc[valid_index] #validation target(Binary for TTF<0.5 Secs)\n\n        #apply min max scaler on training, validation data\n        train_x,valid_x,test_scaled=normalize(train_x.copy(), valid_x.copy(), test.copy(), 'min_max', [])\n\n        #Reshape training,validation,test data for fitting\n        train_x=train_x.values.reshape(train_x.shape[0],1,train_x.shape[1])\n        valid_x=valid_x.values.reshape(valid_x.shape[0],1,valid_x.shape[1])\n        test_scaled=test_scaled.values.reshape(test_scaled.shape[0],1,test_scaled.shape[1])\n\n        #obtain Neural Network Instance\n        #model = get_model()\n        model = get_model_2()#conv1d-conv2d\n\n        #setup Neural Network callbacks\n        cb_checkpoint = ModelCheckpoint(\"model.hdf5\", monitor='val_regressor_mean_absolute_error', save_weights_only=True,save_best_only=True, period=1)\n        cb_Early_Stop=EarlyStopping( monitor='val_regressor_mean_absolute_error',patience=20)\n        cb_Reduce_LR = ReduceLROnPlateau(monitor='val_regressor_mean_absolute_error', factor=0.5, patience=5, verbose=0, mode='auto', min_delta=0.0001, cooldown=0, min_lr=0)\n\n        callbacks = [cb_checkpoint,cb_Early_Stop,cb_Reduce_LR] #define callbacks set\n        \n        ### NN seeds setup- Start\n        os.environ['PYTHONHASHSEED'] = '0'\n        np.random.seed(1234)\n        rn.seed(1234)\n        tf.set_random_seed(1234)\n        session_conf = tf.ConfigProto( allow_soft_placement=True)\n        sess = tf.Session(graph=tf.get_default_graph(), config=session_conf)\n        K.set_session(sess)\n        ### NN seeds setup- End\n        \n        \n        model.fit(train_x,[train_y_ttf,train_y_tsf,train_y_clf],\n                  epochs=1000,callbacks=callbacks\n                  ,batch_size=256,verbose=0,\n                  validation_data=(valid_x,[valid_y_ttf,valid_y_tsf,valid_y_clf]))\n\n        model.load_weights(\"model.hdf5\")\n        \n        oof[valid_index] += model.predict(valid_x)[0].ravel()\n        prediction += model.predict(test_scaled)[0].ravel()/n_fold\n        \n        K.clear_session()\n        \n    # Obtain the MAE for this run.\n    model_score=mean_absolute_error(targets['target'], oof)\n    \n    # Sometimes, the NN performs very badly. This happens if the weights are initialized poorly.\n    # If the MAE is < 2, then the model has performed correctly, and we will use it in the average.\n    if model_score < 2:\n        print('MAE: ', model_score,' Averaged')\n        oof_final += oof/n\n        sub_final += prediction/n\n        i+=1 # Increase i, so we know that we completed a successful run.\n        \n    # If the MAE is >= 2, then the NN has performed badly.\n    # We will reject this run in the average.\n    else:\n        print('MAE: ', model_score,' Not Averaged')\n\nprint('\\nMAE for NN: ', mean_absolute_error(targets['target'], oof_final))","metadata":{"execution":{"iopub.status.busy":"2023-03-17T10:34:24.791072Z","iopub.execute_input":"2023-03-17T10:34:24.791375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-success\" style=\"font-size:30px\">\nCombine NN and LGBM using geometric mean\n</div>\n\nThe geometric mean requires taking product of N oofs, and then taking root(N) of the product. \n\nSince we have two oofs, we multiply the two, then take square root.\n\nWe found the geometric mean to perform slightly better than mean and median.","metadata":{}},{"cell_type":"code","source":"# Square root requires non-negative values, so let us force minima to something small.\nMIN_VALUE = 0.1\n\n# Correct LGBM predictions\noof_LGBM[oof_LGBM < MIN_VALUE] = MIN_VALUE\nsub_LGBM[sub_LGBM < MIN_VALUE] = MIN_VALUE\n\n# Correct NN predictions\noof_final[oof_final < MIN_VALUE] = MIN_VALUE\nsub_final[sub_final < MIN_VALUE] = MIN_VALUE","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('MAE for LGBM was: ', mean_absolute_error(targets['target'], oof_LGBM))\nprint('MAE for NN was  : ', mean_absolute_error(targets['target'], oof_final))\n\noof_geomean = (oof_LGBM * oof_final) ** (1/2)\nsub_geomean = (sub_LGBM * sub_final) ** (1/2)\n\nprint('\\nMAE for geometric mean of LGBM and NN was : ', mean_absolute_error(targets['target'], oof_geomean))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save oof and sub as numpy arrays\nnp.save('oof_LGBM.npy', oof_LGBM)\nnp.save('sub_LGBM.npy', sub_LGBM)\nnp.save('oof_NN.npy', oof_final)\nnp.save('sub_NN.npy', sub_final)\nnp.save('oof_geomean.npy', oof_geomean)\nnp.save('sub_geomean.npy', sub_geomean)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save out the geometric mean submission\nsubmission['time_to_failure'] = sub_geomean\nprint(submission.head())\nsubmission.to_csv('submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save solution using just LGBM\nsubmission['time_to_failure'] = sub_LGBM\nprint(submission.head())\nsubmission.to_csv('submission_LGBM.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save solution using just NN\nsubmission['time_to_failure'] = sub_final\nprint(submission.head())\nsubmission.to_csv('submission_NN.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}