09 - Convolutional Neural Networks (CNN)

Class: CSCE-421


Notes:

Convolutions

Fully Connected Layer

Pasted image 20260219105125.png500
input (1×3072) -> Wx (10×3072) -> activation (1×10)

1 number: the result of taking a dot product between a row of W and the input (a 3072-dimensional dot product)

Convolution Layer

Pasted image 20260219100750.png500

Pasted image 20260219100813.png500

wTx+b

Notes:

Pasted image 20260219101103.png500

Notes:

Pasted image 20260219101206.png500

Notes:

Pasted image 20260219101346.png500

Notes:

Pasted image 20260219101455.png500

Pasted image 20260219101509.png500

f[x,y]∗g[x,y]=∑n1=−∞∞∑n2=−∞∞f[n1,n2]⋅g[x−n1,y−n2]

- elementwise multiplication and sum of a filter and the signal (image)

Notes:

Pasted image 20260219101615.png500

Notes:

Checkpoint 1 (convolution)

1. The Fully Connected Layer (The Old Way) Imagine you have a tiny color image that is 32 pixels wide, 32 pixels high, and has 3 color channels (Red, Green, Blue). In total, that is 32×32×3=3072 numbers. In a Fully Connected (FC) Layer, the network takes all 3072 numbers and stretches them out into one massive, flat line. To generate just one output number, the network uses a filter that is exactly the same size: 3072 weights. It multiplies every single pixel by a corresponding weight and adds them all together (a 3072-dimensional dot product).

2. The Convolution Layer (The New Way) Instead of looking at the whole image at once, a Convolutional Layer looks at small, local patches.

3. Activation Maps (Output Slices) Next, you slide this filter across the image, doing this 75-dimensional dot product at every single spatial location. When you finish sliding the filter across the entire image, you will have created a brand new, flat 2D grid of numbers. This new grid is called an Activation Map (or output slice).

4. Multiple Filters = Multiple Slices One filter can only look for one specific type of feature. But to understand an image, we need to find horizontal edges, diagonal edges, green spots, etc.

5. The Output Size Calculation (Basic) Your notes introduce the fundamental formula for figuring out the size of your new activation map: size of input - size of filter + 1.

A closer look at spatial dimensions:

Pasted image 20260219102031.png450
Pasted image 20260219102110.png450
Pasted image 20260219102154.png450
Pasted image 20260219102207.png450
Pasted image 20260219102224.png450

Notes:

size=input size - filter size2+1

Pasted image 20260219102455.png450
Pasted image 20260219102509.png450
Pasted image 20260219102525.png450
Pasted image 20260219102544.png475
Pasted image 20260219102612.png500

 Output size: (N−F)/ stride +1 e.g. N=7,F=3: stride 1⇒(7−3)/1+1=5 stride 2⇒(7−3)/2+1=3 stride 3⇒(7−3)/3+1=2.33

Notes:


Stride

In practice: Common to zero pad the border

Pasted image 20260219102912.png500

Notes:

Pasted image 20260219103225.png500
e.g. input 7×7
3×3 filter, applied with stride 1
pad with 1 pixel border => what is the output?

Common to see CONV layers with stride 1, filters of size FxF, and zero-padding with (F-1)/2. (will preserve size spatially)

Notes:

Pasted image 20260219103346.png500

Checkpoint 2 (spatial dimensions)

To understand this section, we need to address two major problems that happen when we slide a filter over an image: the shrinking problem and the border problem.

1. The Shrinking Problem When you slide a filter (like a 3×3 box) across an image, the filter must stay completely inside the image boundaries. Because the filter cannot hang off the edge, you cannot place the center of the filter over the very first pixel. As a result, your output is always smaller than your input. If you have a 7×7 image and a 3×3 filter, sliding it one pixel at a time will only yield 5 valid placements, giving you a 5×5 output. If you build a "Deep" Neural Network and apply this convolution repeatedly (e.g., 32→28→24→20), your image will rapidly shrink down to nothing before the network has a chance to learn complex patterns.

2. What is Stride? "Stride" is the step size of your sliding window.

3. Calculating Spatial Dimensions (The Math) To figure out exactly how big your output image will be, your notes provide a formula: Output Size = Input−FilterStride+1. Let's apply this to a 7×7 image with a 3×3 filter:

4. The Border Problem Think about a pixel in the exact center of your image. As the 3×3 filter slides across, that center pixel gets included in 9 different dot-product calculations. Now think about the pixel in the absolute top-left corner. It only gets included in one calculation (when the filter is in the very first position). Because of this, border pixels are treated "unfairly" and their information is largely ignored compared to center pixels.

5. The Solution: Zero Padding To fix both the shrinking problem and the border problem, we use Zero Padding. This means we artificially glue a border of dummy pixels (usually with a value of 0) all the way around the outside of the original image.

If you take a 7×7 image, pad it by 1 (making it 9×9), and run a 3×3 filter over it with stride 1, the output will be exactly 7×7. Padding does not negatively affect the output; it simply allows you to build much deeper networks without your data shrinking out of existence.

Convolution: translation-equivariance

Pasted image 20260219103443.png400

Pasted image 20260219103600.png500

Convolution = local connection + weight-sharing

Pasted image 20260219103705.png500

Notes:

Convolution: linear transform

Pasted image 20260219104430.png500

Notes:

Receptive Field

Pasted image 20260219105052.png500

Notes:

Checkpoint 3 (CNNs & Invariance)

To truly understand Convolutional Neural Networks (CNNs), we need to look at why they work so much better than traditional networks for images. This section covers the fundamental rules and "biases" we build into these models.

1. Translation Equivariance vs. Translation Invariance These two terms sound almost identical but mean very different things in Machine Learning (and professors love to test you on the difference!):

2. How CNNs achieve Invariance (Global Pooling) If convolutions only give us equivariance (the output map shifts), how do we achieve invariance? We use a Global Pooling layer at the very end of the network. Imagine your final 2D feature map detects "cat-ness." If the cat is in the top left, the high numbers are in the top left. If you apply a "Global Max Pooling" operation, you simply ask the computer to find the single highest number in that entire 2D grid and throw the rest of the grid away. Whether the high number was in the top-left or bottom-right, the final extracted number is the same. Location is discarded, and translation invariance is achieved!

3. Convolution is just a restricted Fully Connected (FC) Layer Your notes make a brilliant mathematical connection: a convolution is actually just a Fully Connected layer with two strict rules applied to it.

4. Inductive Bias By applying the two rules above, we are giving the network an Inductive Bias. A neural network is essentially a blank slate. If you give an FC network an image, it has no idea that pixels close together form shapes. It has to learn that from scratch, which takes massive amounts of data. An Inductive Bias is our way of giving the network a hint. By forcing it to use small sliding windows with shared weights, we are embedding our human assumption that "local pixels form features, and features are useful no matter where they appear in the image."

5. The Receptive Field The "receptive field" is simply the area of the original input image that a specific hidden unit is "looking at."

Examples time:

Input volume: 32×32×3
105×5 filters with stride 1 , pad 2

Pasted image 20260224100140.png200

Output volume size: ?

Number of parameters in this layer?

1x1 convolution layers make perfect sense

"it is just used to reduce the number of feature maps -> and therefore reduce the number of parameters"

Pasted image 20260224101911.png425

Notes:

Pooling Layer

Pasted image 20260224102723.png300

MAX POOLING

Pasted image 20260224102749.png400

Notes:

Checkpoint 4 (Examples, 1x1, Pooling layer)

To master Convolutional Neural Networks (CNNs) for your exam, you need to be able to calculate exactly how data changes as it flows through the network. This section breaks down the math, followed by two special techniques used to shrink the data.

1. The Output Size (Feature Map Size) Imagine you are given a 32×32×3 color image (32 width, 32 height, 3 color channels). You want to apply 10 different 5×5 filters to it, using a stride of 1 and a padding of 2. How big is the output?

2. Calculating the Number of Parameters This is a guaranteed exam question! Your professor wants to know if you understand what a model is actually "learning."

3. The Magic of 1x1 Convolutions A 1×1 filter sounds completely useless at first—if it's only looking at 1 pixel at a time, how can it detect an edge or a pattern? The secret is that it looks through the depth. If you have an input that is 56×56×64, a 1×1 filter is actually a 1×1×64 column. It does a 64-dimensional dot product, essentially acting as a mini fully connected layer across the channels for every single pixel.

4. The Pooling Layer (Max Pooling) As you build a deep network, your representations need to become smaller and more manageable. While convolutions extract features, Pooling Layers simply shrink the spatial dimensions (width and height).

Fully Connected Layer

How to perform BP in convolution, pooling layers?

Pasted image 20260224103244.png500

Notes:

Optional for CSCE-421

Backpropagation in Convolutional Neural Networks

Backpropagating through Convolutions

Backpropagation with an Inverted Filter (Single Channel)

a b c
d e f
g h i
Filter during convolution
i h g
f e d
c b a
Filter durin backpropagation

Notes:

Mini-batch SGD

Loop:

  1. Sample a batch of data
  2. Forward prop it through the graph (network), get loss
  3. Backprop to calculate the gradients
  4. Update the parameters using the gradient

Activation Functions

tanh(x)

ReLU (Rectified Linear Unit)

Notes:

TLDR: in practice:

Checkpoint 5 (activation functions)

To understand this section, we first need to recall why we use activation functions. If you stack multiple layers of neurons together but only use linear math (like simple multiplication and addition), the entire network collapses mathematically into just one single linear layer. To allow the network to learn complex, curvy, non-linear shapes (like identifying a cat), we must pass the output of every neuron through a non-linear "activation function" before sending it to the next layer.

Here is the breakdown of the specific activation functions your professor discussed:

1. tanh(x) (Hyperbolic Tangent)

2. ReLU (Rectified Linear Unit) Currently, ReLU is the undisputed king of deep learning activation functions.

3. Leaky ReLU

Data Preprocessing

Learning Rate in Gradient Descent

W:=W−η∂ε∂Wη : learning rate 

Pasted image 20260226093723.png350

Notes:

Normalization

Pasted image 20260226093809.png350

Pasted image 20260226093843.png200

Notes:

Data normalization in machine learning

Pasted image 20260226093919.png600

Notes:

TLDR: IN practice for Images: center only

Not common to normalize variance, to do PCA or whitening

Normalization Modules

Pasted image 20260226094159.png500
Pasted image 20260226094247.png500
Pasted image 20260226095826.png500

Notes:

Checkpoint 6 (preprocessing & normalization)

To understand this section, we need to look at how a neural network actually learns (Gradient Descent) and how the scale of our input data can either make this learning process smooth and easy, or chaotic and impossible.

1. The Learning Rate (η) When a network learns, it updates its weights using Gradient Descent. The formula in your notes is: Wnew=Wold−η∂ε∂W

2. The Problem: Un-normalized Data Imagine predicting house prices using a neural network. Your first input feature is "Square Footage" (values around 2,000) and your second feature is "Number of Bedrooms" (values around 3). Because the math inside a neural network is just multiplication, the feature with the massive numbers (2,000) will cause massive gradients, while the feature with small numbers (3) will cause tiny gradients.

3. The Cure: Data Normalization To fix this, we mathematically force all of our input features to be on the exact same scale. For every single feature (column) in your dataset, you do two things:

4. The Golden Rule of Test Data Your notes highlight a very common, fatal mistake students make. When you normalize your Training Data, you calculate its specific Mean and Standard Deviation. When it is time to evaluate your Test Data, you must use the exact same Mean and Standard Deviation numbers you calculated from the Training Data.

5. TL;DR for Images Images are a special case. Every single pixel in an image is already on the exact same scale: a brightness value between 0 and 255. Because they are naturally on the same scale, dividing by the standard deviation (variance) is unnecessary and mathematically wasteful. For images, we only center the data. We find the "average image" (or the average red, green, and blue values) across the training set, and just subtract that from every image.

6. Normalization Modules (Inside the Network) Data normalization fixes the inputs before they enter the network. But what happens inside? As data passes through layer 1, gets multiplied by weights, and passed through ReLU, the nice, neat scale gets ruined. By the time it reaches layer 10, the numbers might be massively distorted again. To fix this, we insert Normalization Modules (like Batch Normalization) between the hidden layers of the network. This ensures that the data is continuously re-normalized before it enters the next layer, keeping the variance perfectly maintained all the way through the deep network.

Batch Normalization

Batch Normalization

ONE OF THE MOST IMPORTANT TOPICS IN DEEP LEARNING

"you want zero-mean unit-variance activations? just make them so."
consider a batch of activations at some layer. To make each dimension zero-mean unit-variance, apply:

x^(k)=x(k)−E[x(k)]Var[x(k)]

this is a vanilla differentiable function...

Notes:

"you want zero-mean unit-variance activations? just make them so."

  1. compute the empirical mean and variance independently for each dimension.

    Pasted image 20260226100417.png233

    • Fully connected layer with D units
    • For each sample you will get an ... dimensional vector
    • Get the mean, get the standard deviation, subtract mean and divide standard deviation element-wise
  2. Normalize

x^(k)=x(k)−E[x(k)]Var[x(k)]

Notes:

Pasted image 20260226100614.png500

Notes:

Batch Normalization (Math)

Normalize:

x^(k)=x(k)−E[x(k)]Var[x(k)]

And then allow the network to squash the range if it wants to:

y(k)=γ(k)x^(k)+β(k)

Note, the network can learn:

γ(k)=Var[x(k)]β(k)=E[x(k)]

to recover the identity mapping.

Notes:

Checkpoint 7 (Batch Normalization I)

To understand Batch Normalization (often called BatchNorm or BN), we need to remember the core problem with deep neural networks: as data passes through many layers, getting multiplied by weights and passed through activation functions, the scale of the numbers can go completely out of control.

We already solved this for the input data by normalizing it before it enters the network. But what about the data inside the hidden layers? By layer 50, the numbers could be massively shifted, causing gradients to vanish or explode. Batch Normalization is the ingenious solution to this: we simply force the data to be normalized inside the hidden layers too.

Here is the step-by-step breakdown of how it works and the math behind it.

1. The Context: Mini-Batches and Epochs When training deep learning models, we use Mini-batch Stochastic Gradient Descent (SGD).

2. The Core Math: Forcing Normalization Imagine you have a Fully Connected (FC) layer. For a mini-batch of 32 images, a specific hidden node will output 32 different numbers. Batch Normalization says: "Let's take those 32 numbers and normalize them so they have a mean of 0 and a variance of 1."

3. Where does it go? Conceptually, Batch Normalization is treated as its own distinct layer. The standard order of operations in modern neural networks is:

  1. Linear transformation (Fully Connected or Convolutional layer)
  2. Batch Normalization layer
  3. Non-linear Activation Function (like ReLU)

4. The Twist: Scale (γ) and Shift (β) If we rigorously force every single hidden layer to always have a mean of 0 and a variance of 1, we actually restrict what the network can learn. Sometimes, the network wants the data to be spread out or shifted to take full advantage of the activation function!

To fix this, the inventors of BatchNorm added a brilliant mathematical trick: Scale and Shift. After calculating the perfectly normalized data x^, we multiply it by a new parameter γ (Gamma) and add a new parameter β (Beta).

Why is this so important? γ and β are learnable parameters, exactly like the weights in a normal layer. The network uses gradient descent to learn the best possible scale and shift for the data.

(Note: Everything we just discussed happens during Training. Test time behaves differently, which you will see in your next slides!)

Batch Normalization (Algorithm)

 Input: Values of x over a mini-batch: B={x1…m};  Parameters to be learned: γ,β Output: {yi=BNγ,β(xi)}μB←1m∑i=1mxi// mini-batch mean σB2←1m∑i=1m(xi−μB)2// mini-batch variance x^i←xi−μBσB2+ϵ // normalize yi←γx^i+β≡BNγ,β(xi) // scale and shift 

Note: at test time BatchNorm layer functions differently:

Notes:

...

Batch Normalization: Test Time

Input: x:N×D

μj= (Running) average of values  seen during training 

Learnable params:

γ,β:Dσj2= (Running) average of values  seen during training 

Intermediates:

μ,σ:Dx^:N×Dx^i,j=xi,j−μjσj2+ε

Output: y:N×D

yi,j=γjx^i,j+βj

Batch Normalization for ConvNets

Batch Normalization for fully-connected networks

x:N×D Normalize ↓μ,a:1×Dγ,β:1×Dy=γ(x−μ)/a+β

Batch Normalization for convolutional networks (Spatial Batchnorm, BatchNorm2D)

x:N×C×H×W Normalize ↓μ,a:1×C×1×1γ,β:1×C×1×1y=γ(x−μ)/a+β

Notes:

Layer Normalization

Batch Normalization for fully-connected networks

x:N×D Normalize ↓μ,a:1×Dγ,β:1×Dy=γ(x−μ)/a+β

Layer Normalization for fully-connected networks Same behavior at train and test! Can be used in recurrent networks

x:N×D Normalize ↓μ,a:N×1γ,β:1×Dy=γ(x−μ)/a+β

Notes:

Checkpoint 8 (Batch Normalization II)

To truly master normalization for your exam, you need to understand that training a model and testing a model are two completely different environments.

1. Batch Normalization at Test Time (The Rules Change) During training, Batch Normalization calculates the mean (μ) and standard deviation (σ) across a mini-batch of data (like 32 images) to normalize the hidden layers.

2. The Limitation of Mini-Batches (Exam Warning!) Your professor explicitly noted this will be on the exam. Why does Batch Normalization limit our mini-batch size?

3. Batch Normalization for Convolutional Networks (Spatial BatchNorm) When we apply Batch Normalization to a Fully Connected network, we have N samples and D neurons. We calculate exactly D means (one for each neuron, averaged across the N samples).

4. Layer Normalization (The Alternative to Batch Norm) Batch Normalization averages across the batch (N). What if we flip the math by 90 degrees and average across the features (D) instead? This is called Layer Normalization.

Early Stopping

Pasted image 20260226103515.png500

Notes:

Regularization: Dropout

In each forward pass, randomly set some neurons to zero Probability of dropping is a hyperparameter; 0.5 is common

Pasted image 20260226103830.png500

Notes:

How can this possibly be a good idea?

Forces the network to have a redundant representation; Prevents co-adaptation of features
Pasted image 20260226104312.png500

Notes:

Another Interpretation

Pasted image 20260226104505.png150

Notes:

Dropout: Test time

Dropout makes our output random!

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image.png237

 Output  (label)  Input  (image) y=fW(x,z) Random  mask 

Want to "average out" the randomness at test-time

y=f(x)=Ez[f(x,z)]=∫p(z)f(x,z)dz

But this integral seems hard ...

Want to approximate the integral

y=f(x)=Ez[f(x,z)]=∫p(z)f(x,z)dz

Consider a single neuron.
At test time we have: E[a]=w1x+w2y
During training we have:

E[a]=14(w1x+w2y)+14(w1x+0y)+14(0x+0y)+14(0x+w2y)=12(w1x+w2y)

At test time, multiply by (1- dropout) probability

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-1.png140

Notes:

Regularization: A common pattern

Training: Add some kind of randomness

y=fW(x,z)

Testing: Average out randomness (sometimes approximate)

y=f(x)=Ez[f(x,z)]=∫p(z)f(x,z)dz

Regularization: Data Augmentation

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-2.png500

Notes:

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-3.png500

Random crops and scales

Training: sample random crops / scales

ResNet:

  1. Pick random L in range [256, 480]
  2. Resize training image, short side =L
  3. Sample random 224×224 patch

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-4.png100

Testing: average a fixed set of crops

ResNet:

  1. Resize image at 5 scales: {224,256,384,480,640}
  2. For each size, use 10224×224 crops: 4 corners + center,+ flips

Regularization: A common pattern (summary)

Training: Add random noise
Testing: Marginalize over the noise

Examples:
Dropout
Batch Normalization
Data Augmentation
DropConnect

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-5.png325x168

Checkpoint 9 (regularization)

To understand these slides, we have to talk about the biggest enemy in Machine Learning: Overfitting. Deep Neural Networks have millions of parameters. If you train them for too long, they will stop learning general patterns (like "what does a cat look like?") and start perfectly memorizing the training data (like "this specific pixel being blue means it's the cat from picture #4").

To fight this, we use Regularization, which is essentially a set of techniques to deliberately handicap the network so it is forced to learn general, robust features rather than memorizing the data.

1. Early Stopping

2. Regularization: Dropout Dropout is one of the most famous and effective regularization techniques in deep learning.

3. Another Interpretation (The Ensemble Method) There is a second, highly mathematical way to look at why Dropout works so well.

4. Dropout at Test Time

5. Regularization: A Common Pattern Your professor highlights a beautiful, unifying theme across all deep learning regularization techniques (like Batch Normalization, Dropout, and Data Augmentation):

6. Regularization: Data Augmentation Finally, how do we make a network invariant to rotations or translations (meaning it recognizes a cat even if it's upside down or in the corner)?

Transfer Learning

"You need a lot of a data if you want to train/use CNNs"

Transfer Learning with CNNs

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-6.png

Notes:

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-7.png

 very similar  dataset  very different  dataset  very little data  Use Linear  Classifier on  top layer  You’re in  trouble... Try  linear classifier  from different  stages  quite a lot of  data  Finetune a  few layers  Finetune a  larger number  of layers 

Checkpoint 10 (transfer learning)

To understand Transfer Learning, you have to remember one fundamental rule about Deep Learning: training a Convolutional Neural Network (CNN) from scratch requires a massive amount of data. If you only have a small dataset—for example, just 2,000 images of specific dog breeds—trying to train a deep network from scratch will usually result in severe overfitting.

Transfer Learning is the ultimate shortcut to solving this problem.

By doing this, you are effectively transferring the "vision" capabilities the network already learned and simply teaching its final decision-making layers to recognize your specific dog breeds!

Summary:

CNN Architectures

Review: LeNet-5

This network will be similar to what will be on the final
00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-8.png518

Notes:

Case Study: AlexNet

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-9.png

Details/Retrospectives:

Notes:

ImageNet Large Scale Visual Recognition Challenge (ILSVRC) winners

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-10.png700

Notes:

Case Study: VGGNet

Small filters, Deeper networks

8 layers (AlexNet)
-> 16-19 layers (VGG16Net)

Only 3×3 CONV stride 1, pad 1 and 2×2 MAX POOL stride 2

11.7% top 5 error in ILSVRC'13 (ZFNet)
-> 7.3% top 5 error in ILSVRC'14

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-11.png375

Notes:


Q : Why use smaller filters? ( 3×3 conv)

Stack of three 3×3 conv (stride 1) layers has same effective receptive field as one 7×7 conv layer [7x7]

But deeper, more non-linearities

And fewer parameters: 3 * ( 32C2 ) vs. 72C2 for C channels per layer

Notes:


Example:
00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-12.png

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-13.png

Case Study: GoogLeNet

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-14.png304

Apply parallel filter operations on the input from previous layer:

Concatenate all filter outputs together depth-wise

Q: What is the problem with this? [Hint: Computational complexity]

Notes:

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-15.png

Stack Inception modules with dimension reduction on top of each other

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-16.png400

Checkpoint 11 (case studies)

Review: LeNet-5 This is a foundational 1990s architecture consisting of a simple, alternating sequence of convolutional and pooling layers that ends with a fully connected layer.

Case Study: AlexNet This architecture proved that training much larger models is possible by utilizing GPUs, and it was the first major network to successfully introduce the ReLU activation function.

ImageNet The ImageNet visual recognition challenge was the historical turning point that proved to the world that deep learning could vastly outperform traditional machine learning models.

Case Study: VGGNet This model demonstrated that stacking multiple small 3×3 convolutional layers on top of each other gives the exact same "receptive field" as using a single large filter (like 7×7), but it is far more computationally efficient and uses fewer parameters.

Case Study: GoogLeNet This architecture pioneered the "bottleneck" concept, using 1×1 convolutions to drastically shrink the number of feature maps (and thus reduce parameters) before applying more expensive parallel filter operations.

Case Study: ResNet

What happens when we continue stacking deeper layers on a "plain" convolutional neural network?

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-17.png500

56-layer model performs worse on both training and test error
-> The deeper model performs worse, but it's not caused by overfitting!

Notes:


Hypothesis: the problem is an optimization problem, deeper models are harder to optimize

The deeper model should be able to perform at least as well as the shallower model.

A solution by construction is copying the learned layers from the shallower model and setting additional layers to identity mapping.


Solution: Use network layers to fit a residual mapping instead of directly trying to fit a desired underlying mapping

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-19.png417

Notes:


Full ResNet architecture:

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-20.png346

Notes:

"Bottleneck"

For deeper networks (ResNet-50+), use “bottleneck” layer to improve efficiency (similar to GoogLeNet)

image-25.png345x252

Notes:

Identity Mappings in Residual Learning

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-21.png224

Notes:

Identity Mappings in Deep Residual Networks

Improving ResNets...

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-22.png333

There is one more thing that is important:

Notes:

Things to takeaway:

  1. What is the skip operation (ResNet)
  2. Bottleneck
  3. Identity mapping

Checkpoint 12 (ResNet)

To understand ResNet (Residual Networks), we need to look at the biggest mystery that haunted deep learning before 2015: why do deeper networks perform worse?

1. Case Study: ResNet Intuitively, a 56-layer network should be much smarter than a 20-layer network. At the very least, it could just copy the 20 layers and do nothing for the remaining 36 layers. However, researchers found that 56-layer "plain" networks performed substantially worse on both training and test data. This was not caused by overfitting, but by an optimization problem: it is incredibly difficult for Gradient Descent to update weights effectively across 56 layers without the signal vanishing.

The solution was the Residual Block. Instead of forcing a layer to learn a completely new, complex mapping (H(x)), the network forces the layer to learn a "residual" or a difference (F(x)=H(x)−x).

2. "Bottleneck" In extremely deep networks (like ResNet-50 or ResNet-152), applying 3×3 convolutions across hundreds of feature maps becomes far too computationally expensive. To fix this, ResNet uses a "Bottleneck" design (borrowed from GoogLeNet). Instead of one large convolution, the block uses three steps:

  1. A 1×1 convolution to compress/reduce the number of feature maps (the bottleneck).
  2. A standard 3×3 convolution on this smaller, manageable amount of data.
  3. Another 1×1 convolution to expand the feature maps back to their original depth. This drastically reduces the number of parameters while maintaining the exact same network capability.

3. Identity Mappings in Residual Learning For the skip connection (x) and the main convolutional path (F(x)) to be added together (F(x)+x), there is a strict mathematical requirement: they must have the exact same spatial size (width/height) and depth (number of feature maps). If the main path uses a stride of 2 (which shrinks the image) or increases the number of filters, the matrix addition will crash because the dimensions no longer match.

4. Identity Mappings in Deep Residual Networks This refers to an ongoing debate and improvement by the creators of ResNet regarding where to put the activation functions (ReLU) and Batch Normalization (BN) inside the block.

Question 1: What if shortcut mapping h≠ identity ?

(This is not in exam/not in homework)

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-23.png600

00 - TAMU Brain/6th Semester (Spring 26)/CSCE-421/Ex2/Visual Aids/image-24.png593x328

What this paper propose is to make that skip connection along that identity.

Turns out that there is one person who did something similar but added more operations on every block (making it more complex), and it ended up not working.

Identity Mappings in Residual Learning

image-26.png453x199