Caffe Prototxt層系列：Convolution Layer

阿新 • • 發佈：2019-01-14

Convolution Layer是CNN中最常見最重要的特徵提取層，形式多種多樣

首先我們先看一下 InnerProductParameter

message ConvolutionParameter {
	  optional uint32 num_output = 1; // The number of outputs for the layer  輸出特徵圖個數（常以特徵圖為單位）
	  optional bool bias_term = 2 [default = true]; // whether to have bias terms  是否使用bias項
	
	  // Pad, kernel size, and stride are all given as a single value for equal
	  // dimensions in all spatial dimensions, or once per spatial dimension.
	  repeated uint32 pad = 3; // The padding size; defaults to 0   填充大小（畫素為單位）
	  repeated uint32 kernel_size = 4; // The kernel size    卷積核大小 
	  repeated uint32 stride = 6; // The stride; defaults to 1   步幅大小
	  // Factor used to dilate the kernel, (implicitly) zero-filling the resulting
	  // holes. (Kernel dilation is sometimes referred to by its use in the
	  // algorithme à trous from Holschneider et al. 1987.)
	  repeated uint32 dilation = 18; // The dilation; defaults to 1       空洞大小  預設1（空洞卷積中常用，如某些語義分割網路）
	
	  // For 2D convolution only, the *_h and *_w versions may also be used to
	  // specify both spatial dimensions.
	  //2D卷積，當高寬不一致時，常常用下列引數
	  optional uint32 pad_h = 9 [default = 0]; // The padding height (2D only)  高度填充大小      
	  optional uint32 pad_w = 10 [default = 0]; // The padding width (2D only) 寬度填充大小
	  optional uint32 kernel_h = 11; // The kernel height (2D only)  卷積核高度
	  optional uint32 kernel_w = 12; // The kernel width (2D only)  卷積核寬度
	  optional uint32 stride_h = 13; // The stride height (2D only)  y軸步幅
	  optional uint32 stride_w = 14; // The stride width (2D only)  x軸步幅
	
	  optional uint32 group = 5 [default = 1]; // The group size for group conv   組卷積 預設1    例項見shufflenet
	
	  optional FillerParameter weight_filler = 7; // The filler for the weight      卷積權重引數
	  optional FillerParameter bias_filler = 8; // The filler for the bias   偏置項引數
	  enum Engine {
	    DEFAULT = 0;
	    CAFFE = 1;
	    CUDNN = 2;
	  }
	  optional Engine engine = 15 [default = DEFAULT];
	
	  // The axis to interpret as "channels" when performing convolution.
	  // Preceding dimensions are treated as independent inputs;
	  // succeeding dimensions are treated as "spatial".
	  // With (N, C, H, W) inputs, and axis == 1 (the default), we perform
	  // N independent 2D convolutions, sliding C-channel (or (C/g)-channels, for
	  // groups g>1) filters across the spatial axes (H, W) of the input.
	  // With (N, C, D, H, W) inputs, and axis == 1, we perform
	  // N independent 3D convolutions, sliding (C/g)-channels
	  // filters across the spatial axes (D, H, W) of the input.
	  optional int32 axis = 16 [default = 1];     卷積軸，預設通道（3D卷積中，有時序軸卷積情況）
	
	  // Whether to force use of the general ND convolution, even if a specific
	  // implementation for blobs of the appropriate number of spatial dimensions
	  // is available. (Currently, there is only a 2D-specific convolution
	  // implementation; for input blobs with num_axes != 2, this option is
	  // ignored and the ND implementation will be used.)
	  //是否強制使用一般的ND卷積，即使對於具有適當空間維數的blob有特定的實現。(目前只有2d特有的卷積實現;對於num_axes != 2的輸入blob，將忽略此選項，並使用ND卷積)
	  optional bool force_nd_im2col = 17 [default = false];
}

卷積形式太多，一時難以收集全，先舉個例子，以後慢慢更新
例如在MobileNet中：

layer {
	  name: "conv6_3/dwise"
	  type: "Convolution"
	  bottom: "conv6_3/expand/bn"
	  top: "conv6_3/dwise"
	  param {
		    lr_mult: 1        
		    decay_mult: 1  
	  }
	  convolution_param {
		    num_output: 960     //輸出個數
		    bias_term: false     //不使用bias項
		    pad: 1   //填充1個畫素
		    kernel_size: 3   //卷積核大小3*3
		    group: 960   //組個數    則一組個數：num_output/ group （必須整除）
		    weight_filler {
		      type: "msra"    
		    }
		    engine: CAFFE
		  }
}

Caffe Prototxt層系列：Convolution Layer

Convolution Layer是CNN中最常見最重要的特徵提取層，形式多種多樣首先我們先看一下 InnerProductParameter message ConvolutionParameter { optional uint32 num_output = 1; /

Caffe Prototxt 啟用層系列：Sigmoid Layer

Sigmoid Layer 是DL中非線性啟用的一種，在深層CNN中，中間層用得比較少，容易造成梯度消失（當然不是絕對不用）；在GAN或一些網路的輸出層常用到首先我們先看一下 SigmoidParameter message SigmoidParameter { enum

Caffe Prototxt 啟用層系列：TanH Layer

TanH Layer 是DL中非線性啟用的一種，在深層CNN中，中間層用得比較少，容易造成梯度消失（當然不是絕對不用）；在GAN或一些網路的輸出層常用到首先我們先看一下 TanHParameter message TanHParameter { enum Engine {

Caffe層系列：Concat Layer

Concat Layer將多個bottom按照需要聯結一個top 一般特點是：多個輸入一個輸出，多個輸入除了axis指定維度外，其他維度要求一致 message ConcatParameter { // The axis along which to concatenate

Caffe層系列：Dropout Layer

Dropout Layer作用是隨機讓網路的某些節點不工作（輸出置零），也不更新權重；是防止模型過擬合的一種有效方法首先我們先看一下 DropoutParameter message DropoutParameter { optional float dropout_ra

Caffe層系列：Softmax Layer

Softmax Layer作用是將分類網路結果概率統計化，常常出現在全連線層後面 CNN分類網路中，一般來說全連線輸出已經可以結束了，但是全連線層的輸出的數字，有大有小有正有負，人看懂不說，關鍵是訓練時，它無法與groundtruth對應（不在同一量級上），所以用Softmax La

Caffe層系列：InnerProduct Layer

InnerProduct Layer是全連線層，CNN中常常出現在分類網路的末尾，將卷積全連線化或分類輸出結果，當然它的用處很多，不只是分類網路中首先我們先看一下 InnerProductParameter message InnerProductParameter {

Caffe層系列：ReLU Layer

ReLU Layer 是DL中非線性啟用的一種，常常在卷積、歸一化層後面（當然這也不是一定的）首先我們先看一下 ReLUParameter // Message that stores parameters used by ReLULayer message ReLUParam

Caffe層系列：BatchNorm Layer

BatchNorm Layer 是對輸入進行歸一化，消除過大噪點，有助於網路收斂首先我們先看一下 BatchNormParameter message BatchNormParameter { // If false, accumulate global mean/var

Caffe層系列：Scale Layer

Scale Layer是輸入進行縮放和平移，常常出現在BatchNorm歸一化後首先我們先看一下 ScaleParameter message ScaleParameter { // The first axis of bottom[0] (the first input

Caffe層系列：Pooling Layer

Pooling Layer 的作用是將bottom進行下采樣，一般特點是：一個輸入一個輸出首先我們先看一下 PoolingParameter message PoolingParameter { enum PoolMethod { //下采樣方式 MAX =

Caffe層系列：Eltwise Layer

Eltwise Layer是對多個bottom進行操作計算並將結果賦值給top，一般特點：多個輸入一個輸出，多個輸入維度要求一致首先看下Eltwise層的引數： message EltwiseParameter { enum EltwiseOp { PROD = 0;

Caffe層系列：Slice Layer

Slice Layer 的作用是將bottom按照需要切分成多個tops，一般特點是：一個輸入多個輸出首先我們先看一下 SliceParameter message SliceParameter { // The axis along which to slice -- m

【6】Caffe學習系列：Blob,Layer and Net以及對應配置檔案的編寫

深度網路(net)是一個組合模型，它由許多相互連線的層（layers)組合而成。Caffe就是組建深度網路的這樣一種工具，它按照一定的策略，一層一層的搭建出自己的模型。它將所有的資訊資料定義為blobs，從而進行便利的操作和通訊。Blob是caffe框架中一種標準的陣列，一種統一的記憶體介面，它詳細

【5】Caffe學習系列：其它常用層及引數

本文講解一些其它的常用層，包括：softmax_loss層，Inner Product層，accuracy層，reshape層和dropout層及其它們的引數配置。 1、softmax-loss softmax-loss層和softmax層計算大致是相同的。softmax是一個分類器，計算的

【4】Caffe學習系列：啟用層（Activiation Layers)及引數

在啟用層中，對輸入資料進行啟用操作（實際上就是一種函式變換），是逐元素進行運算的。從bottom得到一個blob資料輸入，運算後，從top輸入一個blob資料。在運算過程中，沒有改變資料的大小，即輸入和輸出的資料大小是相等的。輸入：n*c*h*w 輸出：n*c*h*w 常用的啟用函式有

【3】Caffe學習系列：視覺層（Vision Layers)及引數

所有的層都具有的引數，如name, type, bottom, top和transform_param. 本文只講解視覺層（Vision Layers)的引數，視覺層包括Convolution, Pooling, Local Response Normalization (LRN),

【2】Caffe學習系列：資料層及引數

要執行caffe，需要先建立一個模型（model)，如比較常用的Lenet,Alex等，而一個模型由多個屋（layer）構成，每一屋又由許多引數組成。所有的引數都定義在caffe.proto這個檔案中。要熟練使用caffe，最重要的就是學會配置檔案（prototxt）的編寫。層有很多種型別，

Caffe學習系列：啟用層（Activiation Layers)及引數

在啟用層中，對輸入資料進行啟用操作（實際上就是一種函式變換），是逐元素進行運算的。從bottom得到一個blob資料輸入，運算後，從top輸入一個blob資料。在運算過程中，沒有改變資料的大小，即輸入和輸出的資料大小是相等的。輸入：n*c*h*w 輸出：n*c*h*w

Caffe學習系列：模型各層資料和引數視覺化

從輸入的結果和圖示來看，最大的概率是7.17785358e-01，屬於第５類（標號從０開始）。與cifar10中的10種類型名稱進行對比： airplane、automobile、bird、cat、deer、dog、frog、horse、ship、truck 根據測試結果，判斷為dog。測試無誤！

Caffe Prototxt層系列：Convolution Layer

相關推薦